Eyes: lids, blinking, line of sight

Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.

Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.

Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.

Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.

The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.

Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.

The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.

Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.

Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.

Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.

41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Your Name 2026-09-24 18:06:04 -04:00
parent c2603102da
commit e497e9ac4d
11 changed files with 1342 additions and 42 deletions

View file

@ -2,7 +2,9 @@
// All policy lives here, never in the renderer. See docs/design.md,
// "The take is the contract".
import { RIGID, LIPS_OUTER, LIPS_INNER, APERTURE, FACE_OVAL, EYE_INNER, subsampleSlots } from './landmarks.js';
import { RIGID, LIPS_OUTER, LIPS_INNER, APERTURE, FACE_OVAL, EYE_INNER,
EYE_R_RING, EYE_L_RING, EYE_R_CORNERS, EYE_L_CORNERS,
EYE_R_LIDS, EYE_L_LIDS, IRIS_A, IRIS_B, subsampleSlots } from './landmarks.js';
import { fitSimilarity, applySimAll, applySim, fitResidual, procrustesMean, smoothTransforms, movingAverage } from './mathutil.js';
// MediaPipe normalises x by image WIDTH and y by image HEIGHT, so its normalised
@ -25,6 +27,13 @@ export function stabilize(dense, smoothRadius, aspect = 1) {
const raw = rigid.map((r) => fitSimilarity(r, ref));
const tfs = smoothTransforms(raw, smoothRadius);
// The refined mesh appends ten iris points to the 468 face points, but a
// plain mesh does not, and synthetic or hand-fed tracks need not. Checked
// rather than assumed: reading past the end would surface as NaN gaze deep
// downstream instead of as "this track carries no iris".
const hasIris = dense.every((f) => f && f.length > IRIS_B[IRIS_B.length - 1]);
const map = (table) => dense.map((f, i) => applySimAll(tfs[i], pick(f, table, aspect)));
return {
ref,
transforms: tfs,
@ -43,9 +52,221 @@ export function stabilize(dense, smoothRadius, aspect = 1) {
const a = applySimAll(tfs[i], pick(f, APERTURE, aspect));
return Math.hypot(a[0].x - a[1].x, a[0].y - a[1].y);
}),
// Eyes. Lid rings are a feature and get traced like the mouth; corners and
// lid centres are the measurement frame; the iris blocks are raw until
// pairIrises decides which is which.
lidR: map(EYE_R_RING), lidL: map(EYE_L_RING),
cornersR: map(EYE_R_CORNERS), cornersL: map(EYE_L_CORNERS),
lidsR: map(EYE_R_LIDS), lidsL: map(EYE_L_LIDS),
irisA: hasIris ? map(IRIS_A) : null,
irisB: hasIris ? map(IRIS_B) : null,
};
}
/* ---------- eyes ---------- */
const mid = (a, b) => ({ x: (a.x + b.x) / 2, y: (a.y + b.y) / 2 });
const dist = (a, b) => Math.hypot(a.x - b.x, a.y - b.y);
// Which iris block belongs to which eye is RESOLVED FROM THE DATA, not declared
// in a table.
//
// The naming in MediaPipe's own material is viewer-relative in some places and
// subject-relative in others, and the two blocks are otherwise
// indistinguishable. Getting it backwards swaps the irises, which looks almost
// right - each eye still has a disc in roughly the right place - so it survives
// a casual eyeball and then reads as a subtly wall-eyed character for the rest
// of the project. Proximity to the eye's corner midpoint settles it in one
// comparison, is impossible to get wrong, and keeps working if the model is
// ever renumbered.
//
// Voted across every frame rather than read off frame zero: one bad detection
// should not decide the whole shot.
export function pairIrises(stab) {
if (!stab.irisA) return null;
let votes = 0;
for (let f = 0; f < stab.irisA.length; f++) {
const cR = mid(stab.cornersR[f][0], stab.cornersR[f][1]);
votes += dist(stab.irisA[f][0], cR) < dist(stab.irisB[f][0], cR) ? 1 : -1;
}
return votes > 0 ? { right: 'irisA', left: 'irisB' }
: { right: 'irisB', left: 'irisA' };
}
// Per-frame eye measurements, in units of eye width. Measurement only - every
// threshold and every stylisation is applied by the callers.
//
// Everything here stays in HEAD-LOCAL space, which is the same space the mouth
// lives in and the same space the registered photo underlay is drawn in. An
// earlier version pinned each eye into a fixed socket fitted to its corners'
// mean over the shot. That does remove the wobble, but it removes too much: the
// residual from out-of-plane rotation is real motion of the eye relative to the
// head, it is still there in the footage, and pinning it away leaves the drawn
// eyes hanging still over a photo whose eyes are moving. The eye has to track
// the face exactly as the mouth does.
//
// The wobble the socket was aimed at is dealt with the way docs/design.md deals
// with it everywhere else - the bounded contour average, the same knob and the
// same radius the mouth uses - and by placing the iris in the frame of the
// ALREADY-SMOOTHED lid ring, so the iris cannot jitter independently of the eye
// it sits in. See buildEyes in app.js.
export function eyeSignals(stab) {
const N = stab.transforms.length;
const pairing = pairIrises(stab);
const openR = [], openL = [], gazeRaw = [], gazeR = [], gazeL = [];
for (let f = 0; f < N; f++) {
const cR = mid(stab.cornersR[f][0], stab.cornersR[f][1]);
const cL = mid(stab.cornersL[f][0], stab.cornersL[f][1]);
const wR = dist(stab.cornersR[f][0], stab.cornersR[f][1]);
const wL = dist(stab.cornersL[f][0], stab.cornersL[f][1]);
// Openness is the lid gap over the CORNER distance. Normalising by the
// corners rather than by anything derived from the lids keeps the
// denominator rigid, so the ratio measures the lid and nothing else, and
// one threshold carries across takes, faces and framings.
openR.push(dist(stab.lidsR[f][0], stab.lidsR[f][1]) / wR);
openL.push(dist(stab.lidsL[f][0], stab.lidsL[f][1]) / wL);
if (!pairing) {
gazeRaw.push({ x: 0, y: 0 }); gazeR.push({ x: 0, y: 0 }); gazeL.push({ x: 0, y: 0 });
continue;
}
const iR = stab[pairing.right][f][0], iL = stab[pairing.left][f][0];
// Gaze is the iris centre relative to the CORNER MIDPOINT, in eye widths -
// a pure offset WITHIN the eye, with the eye's own position divided out, so
// that quantising it quantises the glance and not the head motion carrying
// it.
//
// Measuring against the lid ring's centroid instead would track the lid:
// every blink pulls that centroid down and would fake a glance at the
// floor, on precisely the frames where the eye is most conspicuous. The
// corners are in RIGID, so this origin and this denominator are both immune
// to the performance they are measuring.
const gR = { x: (iR.x - cR.x) / wR, y: (iR.y - cR.y) / wR };
const gL = { x: (iL.x - cL.x) / wL, y: (iL.y - cL.y) / wL };
// ONE gaze for both eyes, and deliberately so. At 320x200 an iris is a
// handful of pixels and its centre comes from five landmarks on an eye
// twenty pixels wide, so the difference between the two measurements is
// noise, not vergence - and independent per-eye noise reads as wall-eyed
// immediately, which is the most expensive artefact on a face. Openness
// stays per-eye, because a wink is real performance and should survive.
gazeRaw.push({ x: (gR.x + gL.x) / 2, y: (gR.y + gL.y) / 2 });
// Kept separately purely as a diagnostic. The two eyes should agree; when
// they disagree in a sustained way rather than frame to frame, that is not
// noise but out-of-plane head rotation biasing the projected iris offset,
// and no 2D measurement can undo it.
gazeR.push(gR); gazeL.push(gL);
}
return { openR, openL, gazeRaw, gazeR, gazeL, hasIris: !!pairing };
}
// Where "not looking anywhere in particular" sits on THIS face. Everything the
// character does is measured as a departure from it, so getting it wrong does
// not bias the gaze slightly - it re-points the whole performance.
//
// `median` is the default and the safe one: the middle of the take, per axis.
// docs/design.md already gives this rule for the anchor fit - the reference is
// the MEAN configuration over the shot, not one frame - and gaze needs it for
// the same reason. The median rather than the mean because a couple of frames
// of hard glance should not drag the rest-point after them.
//
// `neutral` reads the origin off the take's neutral frame instead, which is
// only correct when there genuinely is a held neutral to read. That frame is
// chosen by MINIMUM MOUTH APERTURE, and a closed mouth says nothing whatever
// about where the eyes are pointed - so on footage with no deliberate neutral
// at the top it is an arbitrary frame, and whichever way the performer happened
// to glance on it becomes "straight ahead" for the entire shot. It is kept
// because it is right when the take was shot for this tool, and because being
// able to switch is how you find out that it was not.
export function gazeOrigin(gazeRaw, mode = 'median', neutral = 0, radius = 2) {
if (mode === 'neutral') {
let sx = 0, sy = 0, n = 0;
// A window, not a single frame: one frame of a five-landmark iris centre is
// worth about a pixel of noise, and that pixel would become a permanent
// squint in the output.
for (let f = neutral - radius; f <= neutral + radius; f++) {
const k = Math.min(gazeRaw.length - 1, Math.max(0, f));
sx += gazeRaw[k].x; sy += gazeRaw[k].y; n++;
}
return { x: sx / n, y: sy / n };
}
const mid1 = (vals) => {
const v = vals.slice().sort((a, b) => a - b);
return v.length % 2 ? v[(v.length - 1) / 2]
: (v[v.length / 2 - 1] + v[v.length / 2]) / 2;
};
return { x: mid1(gazeRaw.map((g) => g.x)), y: mid1(gazeRaw.map((g) => g.y)) };
}
// Snap gaze onto a grid, then require a new cell to hold before it takes.
//
// This is the "Primitive - quantised" row of the part table in docs/design.md,
// and it is not a stylisation imposed on the truth: real eyes move in saccades,
// holding a fixation and then jumping. The smooth drift left in the measurement
// is tracker noise plus head-compensation error, so snapping to a grid and
// requiring a dwell removes the noise and recovers the saccade in the same
// operation - the rare case where the aesthetic rule and the physiology agree.
//
// The dwell is what stops a gaze parked on a cell boundary from chattering
// between two cells forever. It is meaningless without a grid, because
// continuous values never repeat, so step 0 short-circuits both.
export function quantizeGaze(gaze, step, dwell) {
if (!(step > 0)) return gaze.map((g) => ({ x: g.x, y: g.y }));
const q = gaze.map((g) => ({
x: Math.round(g.x / step) * step,
y: Math.round(g.y / step) * step,
}));
if (dwell <= 0 || !q.length) return q;
const out = [];
let live = q[0], pend = q[0], run = 0;
for (const g of q) {
if (g.x === pend.x && g.y === pend.y) run++;
else { pend = g; run = 1; }
if (run > dwell && (pend.x !== live.x || pend.y !== live.y)) live = pend;
out.push(live);
}
return out;
}
// Resolve openness into a shut/open decision per frame.
//
// `dwell` is the same guard the teeth get: a lid hovering at the threshold must
// commit before the state changes, so it cannot flicker.
//
// `hold` is the one that is NOT like the teeth, and it is the whole reason
// blinks are worth special-casing. A blink is 100-150ms, which at 12fps is one
// frame and at 24fps is two or three - and a single frame of closed eye reads
// as a dropped frame, not as a blink. Animators draw a blink over two or three
// drawings for exactly that reason. So once the eye shuts it stays shut for
// `hold` frames, which turns an unreadable flicker into a beat.
//
// The hysteresis runs the other way from the teeth: shutting needs a clear
// signal, and once shut the eye is given the benefit of the doubt on reopening,
// because the lid landmarks are least reliable mid-blink.
export function resolveBlink(open, { cut, dwell, hold }) {
const N = open.length;
const shut = new Array(N).fill(false);
let live = false; // current state
let run = 0; // frames the opposing reading has persisted
let held = 0; // frames spent in the current state
for (let f = 0; f < N; f++) {
const reading = live ? open[f] < cut * 1.35 : open[f] < cut;
if (reading === live) run = 0;
else {
run++;
// Leaving a blink additionally requires the blink to have been on screen
// long enough to be legible; entering one never waits.
if (run > dwell && (!live || held >= hold)) { live = reading; held = 0; run = 0; }
}
held++;
shut[f] = live;
}
return shut;
}
// Stage 4: fixed-index subsample of a stabilised ring, then map from normalised
// face space into character raster space.
export function toRasterRing(stabRing, ringTable, n, xform) {
@ -178,6 +399,24 @@ export function heldFrame(kept, f) {
return hit;
}
// Hold every output frame back onto an exposure grid: 1 = on 1s, 2 = on 2s, and
// so on. Frame 5 at exposure 2 reads the pose from frame 4.
//
// This is where "aesthetic sparseness" belongs. docs/design.md used to put it at
// the extraction rate - pick 12fps and the timing is already chosen - but that
// makes the timing a property of a directory of PNGs, so auditioning 12 against
// 24 means re-ripping the clip and re-running detection over all of it. Rip
// dense once and quantise here instead: the dense track stays at the camera's
// rate, the decision stays reversible, and the audio clock is untouched, so
// sync cannot drift while you try timings.
//
// Floor, never round. Rounding would let an output frame read a pose from the
// FUTURE, which is a lead - a separate control, applied after this one, for a
// separate reason.
export function exposeIndex(f, exposure) {
return exposure > 1 ? Math.floor(f / exposure) * exposure : f;
}
// Shift a performance track against the clock, clamped at the ends.
//
// Pure and exported so the shift can actually be asserted: "the slider feels