diff --git a/README.md b/README.md index 9f1503f..4cb8dda 100644 --- a/README.md +++ b/README.md @@ -43,6 +43,13 @@ playback speed — neither gives a deterministic per-frame pass. assuming, because a guessed fps desynchronises audio from picture — and sync is the one thing this view exists to show. +**Exposure** decides how often the picture gets a new drawing: rip at 24 and +render `on 2s` for 12, `on 3s` for 8. The dense track and the audio are +untouched, so it is a dropdown rather than a re-rip, and the export emits keys +only on the grid instead of the same pose twice. Everything rides the same grid +— mouth, eyes, teeth, plate — because a head cutting on the odd frames while the +mouth cuts on the even ones reads as two performances laid over each other. + **Audio is the playback clock**: `frame = floor(audio.currentTime * fps)`. A slow render loop therefore drops frames instead of drifting, and ½x / ¼x work by setting `playbackRate` with the picture following for free. @@ -62,6 +69,7 @@ is hand-drawn head plates, which this tool does not yet do. | Knob | What it does | | --- | --- | | vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. | +| exposure | How often the picture changes: on 1s, 2s, 3s, 4s. Rip dense, choose timing here. | | mouth lead ±f | Shifts the performance tracks earlier against the audio and the head. `[` `]`. | | contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. | | anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform*. | @@ -99,6 +107,100 @@ kept pixels green, extracted contour amber. Tune against that, not the numbers. brightness and biased low rather than high. Not implemented: it is not visible in the test footage, which reads as a dark cavity with a bright upper-teeth band. +## Eyes + +Three parts per eye, stacked the way the mouth is: a dark **lash ring**, the +**sclera** inside it, and the **iris** inside that, with a square **pupil** in +the iris. The dark ring outside a pale interior is what makes a flat shape read +as an opening rather than a blob, and it is why a blink costs nothing — when the +lid shuts, the traced ring goes flat and the lash line collapses to a lens, +which is a closed eye, drawn correctly, for free. + +The **lids are a feature**, rotoscoped like the mouth: head-local, a key on every +frame, the same `contour avg` knob. They track the face, because the face is what +they are attached to. + +The **iris is a primitive** — a disc at a quantised position — and that is where +the stylisation is. + +### Line of sight + +Gaze is the iris centre relative to the **midpoint of the eye's two corners**, +in units of corner distance. Both corners are in `RIGID`, which is the point: +the origin and the scale are built only from landmarks that do not move under +performance. Measure against the lid ring's centroid instead and every blink +drags that centroid down and fakes a glance at the floor, on exactly the frames +where the eye is most conspicuous. + +**Both eyes share one gaze.** At 320×200 an iris is a handful of pixels and its +centre comes from five landmarks on an eye twenty pixels wide, so the difference +between the two measurements is noise, not vergence — and independent per-eye +noise reads as wall-eyed immediately, which is the most expensive artefact on a +face. Openness stays per-eye, so a wink survives. + +Then the gaze is **quantised to a pixel grid with a dwell**, which is not a +stylisation imposed on the truth: real eyes move in saccades, holding a fixation +and then jumping. The smooth drift left in the measurement is tracker noise plus +head-compensation error, so snapping to a grid and requiring a dwell removes the +noise and recovers the saccade in one operation. The readout reports how many +distinct cells the iris ever occupies — three or four is a character who looks at +things, forty is an unquantised iris sliding around. + +The iris is drawn at the socket read back off the **already-smoothed, already- +subsampled lid ring** — slots 0 and 8 of a 16-slot ring are the corners, and +subsampling to any even budget keeps them at output indices 0 and `n/2`. So the +iris is placed in the frame of the exact polygon it sits inside and cannot drift +relative to its own eye. Size is authored from the take's mean eye width, not +remeasured per frame: a radius that breathes by a fraction of a pixel flickers a +pixel on and off around the whole silhouette. + +The iris is **stencilled to the sclera** and the pupil to the iris — the indexed +buffer is its own clip mask, the way Animator Pro would do it. So the lid crops +the iris at extreme gaze automatically, and nothing needs to clamp the gaze, +which would flatten the performance at exactly the extremes that carry it. + +### Blinking + +Openness is the lid gap over the corner distance — normalised, so one threshold +carries across takes and faces. It gets hysteresis and a dwell like the teeth, +plus one knob the teeth do not have: **blink hold**. A blink is 100–150ms, which +is one frame at 12fps, and a single frame of closed eye reads as a dropped frame +rather than as a blink. Animators draw a blink over two or three drawings for +that reason, so once the eye shuts it stays shut for `hold` frames. + +### The pupil is a square + +At this size a pupil is three pixels across, and a circle of radius 1.5 is not a +circle — it is a plus sign with the corners gnawed off, and it changes shape as +it moves. A square that size is a deliberate mark that stays the same mark +wherever it lands. It is drawn from a rounded centre shared with the iris, so it +is exactly its nominal size on every frame instead of spilling to the next pixel +on some and not others. + +### Which iris is which + +The refined mesh appends ten iris points, five per eye, and MediaPipe's own +left/right naming is viewer-relative in some places and subject-relative in +others. Getting it backwards swaps the irises, which looks *almost* right — each +eye still has a disc roughly where it belongs — so it survives an eyeball and +then reads as a subtly wall-eyed character forever. The pairing is therefore +**resolved from the geometry**, by voting each block's distance to each eye's +corner midpoint across every frame, and the selftest feeds it a track built the +other way round to prove it actually looks. + +| Knob | What it does | +| --- | --- | +| eye vertices | Lid ring vertex budget, off a 16-slot ring. | +| lash line | How far the dark ring sits outside the lid, in pixels. | +| blink cut | Openness below which the eye is shut. Normalised by corner distance. | +| blink hold | Minimum frames a blink stays on screen. A one-frame blink is a dropout. | +| blink dwell | Frames a change must persist. Usually 0 — unlike the teeth, a real blink *is* one frame. | +| gaze gain | Exaggerates or damps the throw. Measured excursion is small; a character usually wants more. | +| gaze step | The pixel grid the iris snaps to. 0 = off, and then dwell does nothing either. | +| gaze dwell | How long a new cell must hold before it takes. Together with step, this is what makes saccades. | +| iris size | Diameter as a percentage of eye width. | +| pupil | Square pupil in whole pixels. 0 = off. | + ## The plate is reference, not art The plate layer has several representations because its job changes. Cycle with @@ -131,6 +233,10 @@ error here. **Aesthetic** sparseness is set by the extraction rate — pick 12fp you have already chosen your timing. **Labour** sparseness is a human drawing each one, and it binds only on the plate. +Aesthetic sparseness is the **exposure** control, not the extraction rate — +making it a render-time grid means auditioning 12 against 24 costs a dropdown +instead of a re-rip and a full re-detection. + So the mouth keeps **every** frame: it is traced, and therefore free. In limited animation lip sync is routinely the densest element, on 1s, while heads hold on 2s and 3s. @@ -162,7 +268,7 @@ chromium --headless --virtual-time-budget=8000 --dump-dom \ http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+' ``` -Or open `selftest.html`. 41 assertions over the stages below detection, plus a +Or open `selftest.html`. 88 assertions over the stages below detection, plus a wiring cross-check: every `el('id')` in `app.js` must exist in `index.html`. A knob wired in one but not the other throws during wiring, which aborts the rest of the module and leaves a blank page — a symptom that points nowhere near its @@ -176,7 +282,7 @@ eyeball. ## Not done yet -Eyes and irises; hand-drawn head plates and per-plate mouth slots (the strip +Brows; hand-drawn head plates and per-plate mouth slots (the strip decides *which frames need one*, but you cannot yet supply the drawing); real performer→character calibration (currently identity, fitting the face oval to the canvas); the override layer; anything on the Animator Pro side. The plate is a diff --git a/docs/design.md b/docs/design.md index a614389..bc7cdbe 100644 --- a/docs/design.md +++ b/docs/design.md @@ -70,9 +70,19 @@ avoidance. ## Two kinds of sparseness -Conflating these was the original design error. **Aesthetic** sparseness is set -by the extraction rate: pick 12fps and the timing is already chosen. **Labour** -sparseness is a human drawing each one, and it binds only on the plate. +Conflating these was the original design error. **Aesthetic** sparseness is the +rate the picture changes at. **Labour** sparseness is a human drawing each one, +and it binds only on the plate. + +Aesthetic sparseness used to be set by the extraction rate — rip at 12 and the +timing is chosen. That was wrong in a small way: it makes the timing a property +of a directory of PNGs, so auditioning 12 against 24 means re-ripping the clip +and re-running detection over the whole of it, and the decision you most want to +play with is the one that costs the most to change. Rip at the camera's rate and +quantise at render time instead — an **exposure** grid, on 1s, 2s, 3s — so the +dense track keeps everything, the audio clock is untouched, and the timing is a +dropdown rather than a re-rip. The take format already carried an `exposure` +field for this; it was simply never driven. So the mouth keeps **every** frame — it is traced, and therefore free. In limited animation lip sync is routinely the densest element, on 1s, while heads hold on @@ -136,6 +146,83 @@ Two escapes, both used: temporal smoothing is well defined, and the star-shaped result suits flat colour. +## Every part is measured in the frame of the thing it is attached to + +The mouth is expressed against the head. The iris is expressed against its own +eye — specifically against the midpoint of that eye's two corners, in units of +corner distance. Both corners are rigid landmarks, so the origin and the scale +of the measurement are immune to the performance being measured. Against the lid +ring's centroid instead, every blink would drag the origin down and fake a glance +at the floor on exactly the frames where the eye is most visible. + +The rule generalises: **measure a feature in a frame built only from landmarks +that do not move with it.** It is the same argument as "rigid landmarks only" for +the anchor fit, one level down. + +There is a tempting over-application. An eye can be pinned into a fixed socket +fitted to its corners' mean over the shot, which removes the residual wobble a +2D similarity cannot — and it is wrong. That residual is real motion of the eye +relative to the head, it is still there in the footage, and removing it leaves +the drawn eyes hanging still over a registered photo whose eyes are moving. A +part must track the face in the same space the underlay is drawn in. The wobble +is a job for the bounded contour average below, not for a second anchor. + +Placement follows from the same idea. The iris is drawn in the frame of the +already-smoothed, already-subsampled lid ring, read off the ring's own corner +vertices, so it cannot drift relative to the eye it sits inside and it inherits +the contour average for free. Size, by contrast, is authored from the take's +mean, never remeasured per frame: a radius that breathes by a fraction of a pixel +flickers a pixel on and off around the whole silhouette. + +## The indexed buffer is its own stencil + +Parts that nest — iris inside sclera, pupil inside iris — clip by colour key: +paint only where the buffer already holds the parent's index. This is how +Animator Pro would do it, it costs one comparison per pixel, and it composes +transitively, so a blink takes the right bite out of the pupil without anything +computing where. + +It also removes a temptation. Without a stencil the gaze has to be clamped to +keep the iris inside the lid, and a clamp flattens the performance at exactly the +extremes that carry it. + +## Quantisation can be the truthful choice + +Gaze snapped to a pixel grid with a dwell is the "Primitive — quantised" row of +the part table, and it looks like a stylisation imposed on a continuous +measurement. It is not. Real eyes move in saccades: hold a fixation, jump, hold. +The smooth drift left in the measured signal is tracker noise plus +head-compensation error. Snapping to a grid and requiring a dwell removes the +noise and recovers the saccade in the same operation — the rare case where the +aesthetic rule and the physiology agree. + +The count of distinct cells the iris ever occupies is the number the knobs exist +to control. Three or four is a character who looks at things; forty is an +unquantised iris sliding around. + +## Some thresholds need a minimum duration, not just a dwell + +A dwell delays a change until it has persisted, which is the right guard against +chatter and is what the teeth use. A blink needs the opposite guard as well. It +lasts 100–150ms — one frame at 12fps — and a single frame of closed eye reads as +a dropped frame rather than as a blink. Animators draw a blink over two or three +drawings for that reason, so once the eye shuts it must stay shut for a minimum +number of frames. Detection accuracy is not the problem; legibility is. + +## Resolve correspondences from data when a wrong guess is survivable + +The refined mesh appends ten iris points, five per eye, and the upstream +left/right naming is viewer-relative in some documentation and subject-relative +in others. Swapping them looks *almost* right — each eye still has a disc roughly +where it belongs — so the error survives inspection and then reads as a subtly +wall-eyed character for the life of the project. + +A hardcoded table is the wrong shape for a fact like that. Voting each block's +distance to each eye's corner midpoint across every frame settles it from the +geometry, cannot be got wrong, and keeps working if the model is renumbered. The +test feeds it a track built the other way round, because a resolver checked only +against the convention it was written for is checking nothing. + ## The bounded smoothing exception *Smooth the transform, never the contour* held while keys were sparse: sampling @@ -188,6 +275,7 @@ Current modules: | Module | Role | | --- | --- | | `landmarks.js` | Index tables. Ring arrays are ordered traversals: slot position *is* vertex identity. | +| `pipeline.js` | …also eye openness, gaze, blink resolution and the iris pairing vote. | | `mathutil.js` | Similarity fit, Procrustes mean, temporal smoothing. | | `pipeline.js` | Stabilise → subsample → key-select → frame-removal. | | `interior.js` | Teeth from image content: Otsu, morphology, components, radial contour. | @@ -202,7 +290,7 @@ Current modules: - **A paint surface.** The plates have nowhere to be drawn. This is the largest gap between "tool" and "suite": a pixel paint canvas with onion skin, palette constraint, and the registered underlay behind it. -- Eyes, irises, brows as parts. +- Brows as parts, and a tongue. - Plate libraries with per-plate mouth slots. - Real performer→character calibration (currently identity). - The override layer. diff --git a/index.html b/index.html index cb59dd6..5902f86 100644 --- a/index.html +++ b/index.html @@ -63,6 +63,21 @@ + + + @@ -86,18 +101,22 @@

source + landmarks

-
outer lip — · inner lip —
+
outer lip — · inner lip — · + lids — · iris —

stabilised (head-local)

-
should sit still except the mouth · grey = held plate outline
dark green ghost = unshifted mouth when lead ≠ 0
+
should sit still except the mouth and eyes · grey = held plate outline
+ dark green ghost = unshifted mouth when lead ≠ 0 · red lid ring = blink

flat render — 320×200 indexed

audio drives the clock — dropped frames, never drift
+ exposure holds the picture on a grid: rip at 24, render on 2s for + 12. The dense track and the audio are untouched, so it is reversible
plate representation: B cycles · photo modes are registered into raster space, so tracing them lands on the mouth
@@ -132,6 +151,16 @@ + + + + + + + + + +
mouth lead shifts the performance tracks earlier (positive) against @@ -150,10 +179,28 @@ brightness; prefer upper biases component choice toward the top of the cavity, where teeth are and the tongue is not. dwell is how many frames a presence change must persist.
+ blink cut is lid gap over corner distance — normalised, so one + value carries across takes. blink hold is the minimum length of a + blink: a real blink is one frame at 12fps and a single frame of closed + eye reads as a dropout, so it is extended to a beat. + pupil is a square, in whole pixels, 0 to turn it off: at this size + a circle of radius 1.5 is a plus sign with the corners gnawed off and it + changes shape as it moves, where a square stays the mark you drew.
+ gaze step is the grid the iris snaps to, in raster pixels, and + gaze dwell is how long a new cell must hold — together they turn + drift into saccades. gaze gain exaggerates or damps the throw; + measured excursion is small and a character usually wants more of it.
suggest tolerance only affects the Suggest button: max head movement allowed before a new drawing is required.
+
+

gaze field

+ +
+
green = every cell the iris visits in the take · + grey = raw · amber = where it is now, quantised
+

teeth measurement

diff --git a/js/app.js b/js/app.js index 752bac5..c52fac1 100644 --- a/js/app.js +++ b/js/app.js @@ -1,10 +1,12 @@ import { FaceLandmarker, FilesetResolver } from 'https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@1.0.1/vision_bundle.mjs'; -import { LIPS_OUTER, LIPS_INNER, FACE_OVAL } from './landmarks.js'; -import { stabilize, toRasterRing, smoothContours, suggestPlateFrames, heldFrame, shiftIndex } from './pipeline.js'; +import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, + EYE_R_RING, EYE_L_RING, IRIS_A, IRIS_B } from './landmarks.js'; +import { stabilize, toRasterRing, smoothContours, suggestPlateFrames, heldFrame, shiftIndex, + exposeIndex, eyeSignals, gazeOrigin, quantizeGaze, resolveBlink } from './pipeline.js'; import { IndexedRaster } from './raster.js'; import { drawRegistered, posterizeInto } from './underlay.js'; import { extractTeeth } from './interior.js'; -import { applySim } from './mathutil.js'; +import { applySim, offsetRing } from './mathutil.js'; import { writeTake } from './take.js'; import { synthDense } from './synth.js'; @@ -16,8 +18,21 @@ const PALETTE = [ { name: 'skin_dark', hex: '#7a4f3a' }, { name: 'mouth_dark', hex: '#24161a' }, { name: 'teeth', hex: '#d9cfc2' }, + // Sclera is not white, and that is authored, not measured. A true white at + // 320x200 next to a warm skin ramp reads as a hole punched in the face; the + // eye sits in a socket, in shadow, so it is a dimmer and cooler tone than the + // teeth, which catch the light. The iris is one dark tone: at this size an + // iris is about five pixels across and a pupil inside it would be one, so the + // iris IS the pupil. Resolving it further would be drawing detail the format + // cannot hold. + { name: 'eye_white', hex: '#c9c3b4' }, + // Three tones for the eye - sclera, iris, pupil - which is the "two or three + // tones per part" budget, spent where it buys the most: an eye with no tonal + // step inside it reads as a hole. + { name: 'iris', hex: '#4a5468' }, + { name: 'pupil', hex: '#171a22' }, ]; -const IDX = { bg: 0, base: 1, dark: 2, mouth: 3, teeth: 4 }; +const IDX = { bg: 0, base: 1, dark: 2, mouth: 3, teeth: 4, white: 5, iris: 6, pupil: 7 }; const state = { dense: null, images: [], stab: null, xform: null, @@ -27,8 +42,11 @@ const state = { fps: 12, audio: null, // fps comes from manifest.json, never guessed aspect: 1, // imgW/imgH; converts MediaPipe's anisotropic space lead: 0, // performance-track offset in frames + exposure: 1, // 1 = on 1s, 2 = on 2s. Picture holds; audio does not. interior: null, // per-frame teeth measurement from image content teeth: null, // resolved per-frame {show, t} after knobs + eyes: null, // resolved per-frame lid rings, shut flags, iris discs + eyeSig: null, // raw eye measurement, kept for the gaze readout }; const el = (id) => { @@ -53,6 +71,19 @@ const opts = () => ({ contourSmooth: +el('contourSmooth').value, apertureThresh: +el('apertureThresh').value / 1000, tol: +el('tol').value / 1000, + exposure: +el('exposure').value, + irisAnchor: el('irisAnchor').value, + gazeOrigin: el('gazeOrigin').value, + eyeVerts: +el('eyeVerts').value, + lashPx: +el('lashPx').value, + irisSize: +el('irisSize').value / 100, + gazeGain: +el('gazeGain').value / 100, + gazeStep: +el('gazeStep').value, // whole raster pixels + pupilPx: +el('pupilPx').value, + gazeDwell: +el('gazeDwell').value, + blinkCut: +el('blinkCut').value / 1000, + blinkHold: +el('blinkHold').value, + blinkDwell: +el('blinkDwell').value, }); function status(msg, kind = '') { @@ -167,6 +198,7 @@ function rebuild(resetKeep) { if (!state.dense) return; const o = opts(); state.lead = o.lead; + state.exposure = o.exposure; const N = state.dense.length; state.stab = stabilize(state.dense, o.smoothWin, state.aspect); @@ -193,6 +225,7 @@ function rebuild(resetKeep) { state.extractKey = extractKey(o); } state.teeth = resolveTeeth(o); + state.eyes = buildEyes(o); // Plate outline per frame, so a kept frame shows its own head shape. state.plates = state.stab.oval.map((r) => r.map(state.xform)); @@ -220,6 +253,124 @@ function faceBoxes() { }); } +// Eyes: lid rings traced per frame, blinks resolved per eye, one gaze shared. +// +// Lids are a FEATURE in the part table - rotoscoped, open vocabulary, a key on +// every frame - so they get exactly the mouth's treatment, including the same +// bounded contour average. The iris is a PRIMITIVE: a disc whose position is +// quantised, which is where the stylisation lives. +function buildEyes(o) { + const st = state.stab, N = state.dense.length; + const sig = eyeSignals(st); + state.eyeSig = sig; + + const blink = { cut: o.blinkCut, dwell: o.blinkDwell, hold: o.blinkHold }; + const shutR = resolveBlink(sig.openR, blink); + const shutL = resolveBlink(sig.openL, blink); + + // Head-local, subsampled, contour-averaged - the identical chain the mouth + // takes, with the identical knob. The eye tracks the face, because the face + // is what it is attached to; what gets removed is per-frame detector jitter, + // not the motion. + const lidR = smoothContours( + st.lidR.map((r) => toRasterRing(r, EYE_R_RING, o.eyeVerts, state.xform)), o.contourSmooth); + const lidL = smoothContours( + st.lidL.map((r) => toRasterRing(r, EYE_L_RING, o.eyeVerts, state.xform)), o.contourSmooth); + + // Where the iris hangs. Three behaviours, because this turns out to be an + // aesthetic choice and not only a correctness one. + // + // STEADY (default) reads the socket back off the DRAWN ring. Slots 0 and 8 of + // a 16-slot lid ring are the two corners, and subsampling to any even budget n + // keeps them at output indices 0 and n/2 - so the ring that gets rendered + // carries its own corners with it. The iris is then placed in the frame of the + // exact polygon it sits inside, after smoothing, after subsampling: it cannot + // drift relative to its own eye, and it inherits the contour average for free. + // + // FREE reads the raw per-frame corners instead, jitter and all. It is what the + // eyes did before any of this, and it is not simply worse - the detector noise + // reads as liveliness, the eye never sits perfectly still, and against flat + // hand-drawn plates that restlessness can be the thing that sells it. It is + // also the honest baseline to compare the other two against. + // + // LOCKED pins the socket to the take's mean, so the eye never moves in the + // head at all. Watch it against a photo underlay and the drawn eyes hang still + // over a face whose eyes are moving - that is the registration cost, and it is + // real - but once the plate is a drawing rather than a photograph, nothing is + // being registered against and it reads as a deliberately locked-off stare. + const ringSocket = (ring) => { + const a = ring[0], b = ring[ring.length / 2]; + return { cx: (a.x + b.x) / 2, cy: (a.y + b.y) / 2, w: Math.hypot(a.x - b.x, a.y - b.y) }; + }; + const rawSocket = (corners, f) => { + const a = state.xform(corners[f][0]), b = state.xform(corners[f][1]); + return { cx: (a.x + b.x) / 2, cy: (a.y + b.y) / 2, w: Math.hypot(a.x - b.x, a.y - b.y) }; + }; + const meanSocket = (rings) => { + const acc = rings.reduce((a, r) => { + const k = ringSocket(r); + return { cx: a.cx + k.cx, cy: a.cy + k.cy, w: a.w + k.w }; + }, { cx: 0, cy: 0, w: 0 }); + const n = rings.length; + return { cx: acc.cx / n, cy: acc.cy / n, w: acc.w / n }; + }; + const socketFor = (rings, corners) => { + if (o.irisAnchor === 'locked') { const k = meanSocket(rings); return () => k; } + if (o.irisAnchor === 'free') return (f) => rawSocket(corners, f); + return (f) => ringSocket(rings[f]); + }; + const skR = socketFor(lidR, st.cornersR), skL = socketFor(lidL, st.cornersL); + const socket = (ring) => ringSocket(ring); + + // Iris radius comes from the take's MEAN eye width, not the current frame's. + // Size is authored; only position is tracked. A radius recomputed per frame + // would breathe by a fraction of a pixel as the fit's depth-scale wanders, + // and at this resolution a fraction of a pixel is a pixel flicking on and off + // around the whole silhouette. + const meanW = (rings) => rings.reduce((a, r) => a + ringSocket(r).w, 0) / rings.length; + const wR = meanW(lidR), wL = meanW(lidL), w = (wR + wL) / 2; + + // Calibrate against the neutral, apply the artist's gain, and only then + // quantise - the grid should be a grid of DRAWN positions, because that is + // what a viewer reads. Gain is an authored parameter: measured gaze excursion + // is small and a character's eye usually wants more throw than a performer's, + // which is a decision for a person and not for the detector. + const origin = gazeOrigin(sig.gazeRaw, o.gazeOrigin, state.neutral); + state.gazeOriginValue = origin; + const px = sig.gazeRaw.map((g) => ({ + x: (g.x - origin.x) * o.gazeGain * w, + y: (g.y - origin.y) * o.gazeGain * w, + })); + const gaze = quantizeGaze(px, o.gazeStep, o.gazeDwell); + + const eye = (sk, lids, shut, rad, f) => { + const e = sk(f); + return { + // The lash line is the lid ring pushed outward by a fixed number of + // pixels, exactly as the mouth's outer ring sits outside its inner one. + // When the eye shuts, the traced ring goes near-degenerate and this + // collapses to a lens - which is a closed eye, drawn correctly, for free. + lash: offsetRing(lids[f], o.lashPx), + lid: lids[f], + shut: shut[f], + // Rounded to whole pixels. The rasteriser quantises everything anyway, so + // this costs nothing - but it means the iris and the square pupil share + // one integer centre, so the pupil is exactly its nominal size on every + // frame instead of spilling to the next pixel on some and not others. + iris: { x: Math.round(e.cx + gaze[f].x), y: Math.round(e.cy + gaze[f].y), r: rad }, + pupil: o.pupilPx, + }; + }; + + return { + gazePx: px, gaze, shutR, shutL, hasIris: sig.hasIris, + frames: Array.from({ length: N }, (_, f) => ({ + r: eye(skR, lidR, shutR, (wR * o.irisSize) / 2, f), + l: eye(skL, lidL, shutL, (wL * o.irisSize) / 2, f), + })), + }; +} + // Presence gets hysteresis and a minimum dwell, the same treatment plate // selection gets: a teeth block that blinks on and off for single frames is // worse than one that is simply absent. Appearing needs a clear signal, staying @@ -275,7 +426,7 @@ function resolveTeeth(o) { // stand-in until a drawing exists. function renderFrame(f, mode = plateMode()) { const r = new IndexedRaster(RW, RH); - const pf = heldFrame(keptSorted(), f); // the plate frame on screen + const pf = plateIndex(f); // the plate frame on screen if (mode === 'posterize' && state.images[pf]) { posterizeInto(r, state.images[pf], state.stab.transforms[pf], state.xform, @@ -285,6 +436,12 @@ function renderFrame(f, mode = plateMode()) { if (mode === 'oval' || mode === 'oval+photo') r.fillPoly(state.plates[pf], IDX.base); } + // Eyes run on the CLOCK, not on the mouth lead. The lead is a lip-sync + // device: it exists because a mouth shape anticipates the sound it makes. + // Nothing about a blink or a glance is tied to the audio, so shifting the + // eyes would only slide them off the head that carries them. + if (state.eyes) drawEyes(r, state.eyes.frames[f]); + const mf = leadIndex(f); // performance frame, possibly ahead r.fillPoly(state.outer[mf], IDX.dark); // mouth keeps every frame if (!state.hidden[mf]) { @@ -295,13 +452,34 @@ function renderFrame(f, mode = plateMode()) { return r; } +// Lash ring, then sclera, then iris - the same three-layer structure the mouth +// has, for the same reason: the dark ring outside the pale interior is what +// makes a flat shape read as an opening rather than a blob. +// +// The iris is stencilled to the sclera it was just drawn over, so the lid crops +// it automatically. Nothing needs to clamp the gaze to keep the iris inside the +// eye, which matters because a clamp would flatten the performance at exactly +// the extremes that carry it. +function drawEyes(r, e) { + for (const s of [e.r, e.l]) { + r.fillPoly(s.lash, IDX.dark); + if (s.shut) continue; // a shut eye IS the lash line, alone + r.fillPoly(s.lid, IDX.white); + r.fillDisc(s.iris.x, s.iris.y, s.iris.r, IDX.iris, IDX.white); + // Stencilled to the iris, which is itself stencilled to the sclera - so the + // pupil is cropped by the lid transitively, and a blink or an extreme gaze + // takes the right bite out of it without anything having to compute where. + if (s.pupil) r.fillRect(s.iris.x, s.iris.y, s.pupil, IDX.pupil, IDX.iris); + } +} + const plateMode = () => el('plateMode').value; // Photo modes composite under the indexed layer, so the flat shapes stay exactly // as they render while the reference sits behind them. function compositeRender(canvas, f, zoom) { const mode = plateMode(); - const pf = heldFrame(keptSorted(), f); + const pf = plateIndex(f); const img = state.images[pf]; const showPhoto = img && (mode === 'photo' || mode === 'photo-dim' || mode === 'oval+photo'); @@ -347,11 +525,22 @@ const keptSorted = () => [...state.keep].sort((a, b) => a - b); // Positive lead = the mouth arrives earlier. Only performance parts shift; the // head stays with the audio, because it is the mouth that should anticipate. function leadIndex(f) { - // Reads a cached scalar, not opts(): this runs once per strip thumbnail, and + // Reads cached scalars, not opts(): this runs once per strip thumbnail, and // calling opts() here meant ~14 DOM reads x 74 frames on every redraw. - return shiftIndex(f, state.lead, state.dense.length); + // + // Exposure first, then lead. The grid decides WHICH frames get a new drawing; + // the lead then shifts which pose that drawing carries, by whole frames of the + // original track. Applying them the other way round would put the changes on + // the wrong beats - the picture would update on the odd frames instead of + // holding on the twos. + return shiftIndex(exposeIndex(f, state.exposure), state.lead, state.dense.length); } +// The plate rides the same grid, so the whole picture updates together. On 2s +// means on 2s - a head that cut on the odd frames while the mouth cut on the +// even ones would read as two performances laid over each other. +const plateIndex = (f) => heldFrame(keptSorted(), exposeIndex(f, state.exposure)); + function blit(canvas, raster, zoom) { canvas.width = RW * zoom; canvas.height = RH * zoom; canvas.getContext('2d').putImageData(raster.toImageData(PALETTE.map((p) => p.hex), zoom), 0, 0); @@ -372,21 +561,42 @@ function drawReadout() { el('readout').textContent = `${state.dense.length} frames → ${kept.length} drawings · ` + `teeth on ${teethFrames}f · ` + + `${blinkRuns(state.eyes.shutR).length}/${blinkRuns(state.eyes.shutL).length} blinks R/L · ` + + `${gazeCells(state.eyes.gaze)} gaze cells · ` + + (state.exposure > 1 + ? `on ${state.exposure}s = ${(state.fps / state.exposure).toFixed(4).replace(/\.?0+$/, '')}fps · ` + : '') + (lead ? `mouth leads ${lead}f (${(lead / state.fps * 1000).toFixed(0)}ms) · ` : '') + `holds ${Math.min(...runs)}–${Math.max(...runs)} frames · ` + `neutral f${state.neutral} · residual ` + `${(state.stab.residual.reduce((a, b) => a + b, 0) / state.dense.length).toFixed(4)}`; } +// Blinks as RUNS, not as shut frames: a three-frame blink is one blink, and the +// count is only useful as "did the performer blink six times or sixty". +function blinkRuns(shut) { + const runs = []; + for (let f = 0; f < shut.length; f++) { + if (shut[f] && !shut[f - 1]) runs.push(f); + } + return runs; +} + +// How many distinct positions the iris ever occupies. This is the number the +// gaze knobs exist to control: two or three is a character who looks at things, +// forty is an unquantised iris sliding around, which is what the grid is for. +const gazeCells = (gaze) => new Set(gaze.map((g) => `${g.x},${g.y}`)).size; + function drawPanes() { const f = state.frame, kept = keptSorted(); - const pf = heldFrame(kept, f); + const pf = plateIndex(f); // The mouth frame is always shown, not only when shifted, so the number can be // watched diverging from f rather than taken on trust. const lead = state.lead; el('framelabel').textContent = `f ${f} / ${state.dense.length - 1} · ${(f / state.fps).toFixed(2)}s · ` + `plate f${pf} · mouth f${leadIndex(f)}` + + (state.exposure > 1 && f % state.exposure ? ' (held)' : '') + (lead ? ` (${lead > 0 ? '+' : ''}${lead} = ${(lead / state.fps * 1000).toFixed(0)}ms)` : '') + (state.keep.has(f) ? ' · KEPT' : ' · held'); @@ -404,12 +614,14 @@ function drawPanes() { const map = (p) => ({ x: dx + (p.x * im.naturalWidth - sx) * s, y: dy + (p.y * im.naturalHeight - sy) * s }); strokePts(g1, LIPS_OUTER.map((i) => map(state.dense[f][i])), '#4ade80'); strokePts(g1, LIPS_INNER.map((i) => map(state.dense[f][i])), '#f87171'); + drawEyeOverlay(g1, map, f); } else { g1.fillStyle = '#555'; g1.font = '13px system-ui'; g1.fillText('synthetic — no source frames', 14, 24); const sc = (p) => ({ x: p.x * c1.width, y: p.y * c1.height }); strokePts(g1, LIPS_OUTER.map((i) => sc(state.dense[f][i])), '#4ade80'); strokePts(g1, LIPS_INNER.map((i) => sc(state.dense[f][i])), '#f87171'); + drawEyeOverlay(g1, sc, f); } const c2 = el('cv-stab'), g2 = c2.getContext('2d'); @@ -426,9 +638,18 @@ function drawPanes() { if (mf !== f) strokePts(g2, z(state.outer[f]), '#2f6b46'); strokePts(g2, z(state.outer[mf]), '#4ade80'); if (!state.hidden[mf]) strokePts(g2, z(state.inner[mf]), '#f87171'); + for (const e of [state.eyes.frames[f].r, state.eyes.frames[f].l]) { + strokePts(g2, z(e.lid), e.shut ? '#f87171' : '#60a5fa'); + if (e.shut) continue; + g2.strokeStyle = '#fbbf24'; + g2.beginPath(); + g2.arc(e.iris.x * ZOOM, e.iris.y * ZOOM, e.iris.r * ZOOM, 0, Math.PI * 2); + g2.stroke(); + } compositeRender(el('cv-render'), f, ZOOM); drawInteriorDebug(f); + drawGazeDebug(f); } // What the teeth measurement actually saw: sampled region, pixels above @@ -458,6 +679,78 @@ function drawInteriorDebug(fRaw) { `area ${m.area}px · ${te.show ? 'SHOWN' : 'hidden'}`; } +// Lid rings and the iris, on the raw frame. Landmark overlays are how you tell +// a tracking failure from a knob set wrong, and the eyes need it more than the +// mouth does: an iris that has latched onto an eyebrow looks, in the flat +// render alone, exactly like a gaze gain that is too high. +function drawEyeOverlay(g, map, f) { + const lm = state.dense[f]; + for (const ring of [EYE_R_RING, EYE_L_RING]) { + strokePts(g, ring.map((i) => map(lm[i])), '#60a5fa'); + } + if (!state.eyes.hasIris) return; + for (const iris of [IRIS_A, IRIS_B]) { + strokePts(g, iris.slice(1).map((i) => map(lm[i])), '#fbbf24'); + } +} + +// The gaze field: every position the iris takes over the whole take, plus where +// it is now. Tune against this, not against the numbers - "4 cells" tells you +// the quantisation is working, but only the picture tells you whether the four +// are the four looks the performance actually has. +function drawGazeDebug(f) { + const cv = el('cv-gaze'), S = 150; + cv.width = S; cv.height = S; + const g = cv.getContext('2d'); + g.fillStyle = '#0d0f16'; g.fillRect(0, 0, S, S); + + const ex = state.eyes; + // Scale so the widest excursion in the take fills the box, with a floor so a + // nearly-still gaze does not get magnified into a light show. + let m = 2; + for (const p of ex.gazePx) m = Math.max(m, Math.abs(p.x), Math.abs(p.y)); + const k = (S / 2 - 8) / m; + const X = (v) => S / 2 + v * k, Y = (v) => S / 2 + v * k; + + const o = opts(); + if (o.gazeStep > 0) { + g.strokeStyle = '#1b2030'; g.lineWidth = 1; + for (let i = -20; i <= 20; i++) { + const v = i * o.gazeStep; + if (Math.abs(v) > m) continue; + g.beginPath(); g.moveTo(X(v), 0); g.lineTo(X(v), S); g.stroke(); + g.beginPath(); g.moveTo(0, Y(v)); g.lineTo(S, Y(v)); g.stroke(); + } + } + g.strokeStyle = '#2a2f3e'; + g.beginPath(); g.moveTo(S / 2, 0); g.lineTo(S / 2, S); + g.moveTo(0, S / 2); g.lineTo(S, S / 2); g.stroke(); + + g.fillStyle = '#2f6b46'; + for (const p of ex.gaze) g.fillRect(X(p.x) - 1.5, Y(p.y) - 1.5, 3, 3); + + const raw = ex.gazePx[f], q = ex.gaze[f]; + g.fillStyle = '#8891a5'; + g.fillRect(X(raw.x) - 1, Y(raw.y) - 1, 2, 2); + g.fillStyle = '#fbbf24'; + g.beginPath(); g.arc(X(q.x), Y(q.y), 4, 0, Math.PI * 2); g.fill(); + + const sig = state.eyeSig, fr = state.eyes.frames[f]; + const og = state.gazeOriginValue; + // Per-eye raw gaze is the diagnostic for a wrong-looking eyeline. If the two + // agree and both point the wrong way, the ORIGIN is wrong. If they disagree in + // a sustained way, it is out-of-plane head rotation biasing the projection, + // which no 2D measurement can undo. + const sgn = (v) => `${v >= 0 ? '+' : ''}${v.toFixed(3)}`; + el('eyeinfo').textContent = + `open R ${sig.openR[f].toFixed(3)} L ${sig.openL[f].toFixed(3)} / cut ${o.blinkCut.toFixed(3)}\n` + + `${fr.r.shut ? 'R SHUT ' : ''}${fr.l.shut ? 'L SHUT' : ''}${!fr.r.shut && !fr.l.shut ? 'both open' : ''}\n` + + `gaze ${q.x >= 0 ? '+' : ''}${q.x.toFixed(1)}, ${q.y >= 0 ? '+' : ''}${q.y.toFixed(1)} px\n` + + `raw R ${sgn(sig.gazeR[f].x)} L ${sgn(sig.gazeL[f].x)} (x, eye widths)\n` + + `origin ${o.gazeOrigin} ${sgn(og.x)}, ${sgn(og.y)}` + + (ex.hasIris ? '' : ' — no iris landmarks'); +} + function strokePts(g, pts, color, lw = 1) { g.strokeStyle = color; g.lineWidth = lw; g.beginPath(); @@ -538,12 +831,57 @@ function drawWorksheet() { /* ---------- export ---------- */ +// Six parts, three per eye, mirroring the mouth's lash/interior/content stack. +// `clip` is what tells the renderer the iris is stencilled by the sclera rather +// than merely drawn after it - without it an extreme gaze would put the iris on +// the cheek. +function eyeParts(grid) { + const out = []; + // Eyes ride the exposure grid but NOT the mouth lead: the lead is a lip-sync + // device and nothing about a blink is tied to the audio. + const src = (f) => exposeIndex(f, state.exposure); + [['r', 20], ['l', 23]].forEach(([side, z]) => { + const at = (f) => state.eyes.frames[src(f)][side]; + out.push( + { name: `eye_${side}`, kind: 'poly', z, color: 'skin_dark', interp: 'hold', + keys: grid.map((f) => ({ f, src: src(f), pts: at(f).lash })) }, + { name: `eye_${side}_in`, kind: 'poly', z: z + 1, color: 'eye_white', interp: 'hold', + parent: `eye_${side}`, + keys: grid.map((f) => + (at(f).shut ? { f, hidden: true } : { f, src: src(f), pts: at(f).lid })) }, + { name: `iris_${side}`, kind: 'disc', z: z + 2, color: 'iris', interp: 'hold', + parent: `eye_${side}_in`, clip: `eye_${side}_in`, + keys: grid.map((f) => { + const e = at(f); + return e.shut ? { f, hidden: true } + : { f, src: src(f), c: { x: e.iris.x, y: e.iris.y }, r: e.iris.r }; + }) }, + ); + if (!state.eyes.frames[0].r.pupil) return; + out.push( + { name: `pupil_${side}`, kind: 'rect', z: z + 3, color: 'pupil', interp: 'hold', + parent: `iris_${side}`, clip: `iris_${side}`, + keys: grid.map((f) => { + const e = at(f); + return e.shut ? { f, hidden: true } + : { f, src: src(f), c: { x: e.iris.x, y: e.iris.y }, size: e.pupil }; + }) }, + ); + }); + return out; +} + function exportTake() { const kept = keptSorted(); const N = state.dense.length; + // Output frames that actually carry a key. Everything between them is a hold, + // which the take format already expresses, so on 2s emits half the keys rather + // than emitting each pose twice. + const grid = []; + for (let f = 0; f < N; f += state.exposure) grid.push(f); const take = { name: el('takename').value || 'line_01', - frames: N, width: RW, height: RH, exposure: 1, fps: state.fps, + frames: N, width: RW, height: RH, exposure: state.exposure, fps: state.fps, palette: PALETTE, slot: { x: RW / 2, y: RH / 2 }, parts: [ @@ -554,14 +892,19 @@ function exportTake() { // key f carries the pose from source frame f+lead - so the renderer never // needs to know about it. { name: 'mouth', kind: 'poly', z: 30, color: 'skin_dark', interp: 'hold', - keys: state.outer.map((_, f) => ({ f, src: leadIndex(f), pts: state.outer[leadIndex(f)] })) }, + keys: grid.map((f) => ({ f, src: leadIndex(f), pts: state.outer[leadIndex(f)] })) }, { name: 'mouth_in', kind: 'poly', z: 31, color: 'mouth_dark', interp: 'hold', parent: 'mouth', - keys: state.inner.map((_, f) => { + keys: grid.map((f) => { const m = leadIndex(f); return state.hidden[m] ? { f, hidden: true } : { f, src: m, pts: state.inner[m] }; }) }, + // Eyes. The lids are traced, so like the mouth they cost nothing and keep + // every frame. The iris is a primitive: its quantised position means the + // key stream is dense but the VALUES change only on saccades, so a + // hold-interpolating renderer cuts between fixations by itself. + ...eyeParts(grid), { name: 'teeth', kind: 'poly', z: 32, color: 'teeth', interp: 'hold', parent: 'mouth_in', - keys: state.teeth.map((_, f) => { + keys: grid.map((f) => { const m = leadIndex(f), te = state.teeth[m]; return te.show && te.pts ? { f, src: m, pts: te.pts } : { f, hidden: true }; }) }, @@ -577,7 +920,9 @@ function exportTake() { a.href = URL.createObjectURL(new Blob([text], { type: 'text/plain' })); a.download = `${take.name}.take`; a.click(); - status(`exported — ${kept.length} plate drawings, ${N} mouth frames`, 'ok'); + status(`exported — ${kept.length} plate drawings, ${grid.length} mouth keys` + + (state.exposure > 1 ? ` on ${state.exposure}s` : '') + ', ' + + `${blinkRuns(state.eyes.shutR).length + blinkRuns(state.eyes.shutL).length} blinks`, 'ok'); } /* ---------- wiring ---------- */ @@ -605,6 +950,7 @@ async function runFrames() { state.interior = measureAll(images, dense, opts()); el('scrub').max = dense.length - 1; state.frame = 0; + labelExposure(); rebuild(true); const dur = (dense.length / state.fps).toFixed(2); status(`${images.length} frames · ${images[0].naturalWidth}x${images[0].naturalHeight} · ` + @@ -626,24 +972,44 @@ function runSynthetic() { state.dense = synthDense(72); el('scrub').max = 71; state.frame = 0; + labelExposure(); rebuild(true); status('synthetic — exercises everything below detection', 'ok'); } +// How each slider's raw value reads out. A table rather than the conditional +// chain this used to be: that chain grew a branch per knob and was one ternary +// away from being unreadable. +const FMT = { + apertureThresh: (v) => (v / 1000).toFixed(3), + tol: (v) => (v / 1000).toFixed(3), + blinkCut: (v) => (v / 1000).toFixed(3), + teethOn: (v) => (v / 100).toFixed(2), + teethErode: (v) => (v / 100).toFixed(2), + tongueReject: (v) => (v / 100).toFixed(2), + topBias: (v) => (v / 100).toFixed(2), + irisSize: (v) => `${v}%`, + gazeGain: (v) => (v / 100).toFixed(2), + gazeStep: (v) => (v ? `${v}px` : 'off'), + pupilPx: (v) => (v ? `${v}px` : 'off'), + lashPx: (v) => `${v}px`, + lead: (v) => (v > 0 ? `+${v}` : String(v)), +}; + for (const id of ['verts', 'smoothWin', 'contourSmooth', 'apertureThresh', 'tol', 'teethOn', 'teethDwell', 'teethErode', 'tongueReject', 'blobGrow', - 'topBias', 'teethVerts', 'teethSmooth', 'lead']) { + 'topBias', 'teethVerts', 'teethSmooth', 'lead', + 'eyeVerts', 'lashPx', 'irisSize', 'pupilPx', 'gazeGain', + 'gazeStep', 'gazeDwell', 'blinkCut', 'blinkHold', 'blinkDwell']) { + const show = () => { + el(id + 'v').textContent = FMT[id] ? FMT[id](+el(id).value) : el(id).value; + }; el(id).addEventListener('input', () => { - el(id + 'v').textContent = id === 'apertureThresh' || id === 'tol' - ? (+el(id).value / 1000).toFixed(3) - : ['teethOn', 'teethErode', 'tongueReject', 'topBias'].includes(id) - ? (+el(id).value / 100).toFixed(2) - : id === 'lead' && +el(id).value > 0 ? `+${el(id).value}` - : el(id).value; + show(); if (id === 'tol') return; // tol only matters when you ask for a suggestion rebuild(false); }); - el(id + 'v').textContent = el(id).value; + show(); } function seekTo(f) { @@ -657,6 +1023,22 @@ el('btn-frames').onclick = runFrames; el('btn-synth').onclick = runSynthetic; el('btn-export').onclick = exportTake; el('plateMode').addEventListener('change', () => { if (state.dense) drawAll(); }); +el('exposure').addEventListener('change', () => { if (state.dense) rebuild(false); }); +for (const id of ['irisAnchor', 'gazeOrigin']) { + el(id).addEventListener('change', () => { if (state.dense) rebuild(false); }); +} + +// Label the exposure options in the only units that mean anything here: the +// rate the picture actually changes at, which depends on the clip's own rate. +// "on 2s" is the animator's name for it and the number is what you hear against +// the audio, so the menu says both. +function labelExposure() { + for (const opt of el('exposure').options) { + const n = +opt.value; + const rate = (state.fps / n).toFixed(4).replace(/\.?0+$/, ''); + opt.textContent = `${rate} fps · on ${n}s`; + } +} el('btn-saveframe').onclick = () => { if (!state.dense) return; const cv = document.createElement('canvas'); @@ -775,3 +1157,5 @@ else if (location.hash === '#frames') runFrames(); else status('ready — Load frames, then step with \u2190 \u2192 and delete with X'); window.__roto = state; // headless smoke test reads this +window.__render = compositeRender; // ...and renders arbitrary frames off-screen +window.__lead = leadIndex; // ...and resolves the performance frame diff --git a/js/landmarks.js b/js/landmarks.js index 8251e6b..ee0094b 100644 --- a/js/landmarks.js +++ b/js/landmarks.js @@ -52,3 +52,46 @@ export function subsampleSlots(len, n) { export function subsampleRing(ring, n) { return subsampleSlots(ring.length, n).map((s) => ring[s]); } + +// ---- eyes ---- +// +// Eyelid rings, under the same contract as the lip rings: ORDERED traversals +// where slot position IS vertex identity. Both eyes start at the OUTER corner +// and go over the UPPER lid first, so slot k means the same anatomy on both +// sides. On a 16-slot ring that puts the four cardinals exactly on the four +// quarter slots - 0 outer corner, 4 upper lid centre, 8 inner corner, 12 lower +// lid centre - so every even vertex budget lands on real landmarks. +// +// The two rings traverse opposite directions on screen, because they are +// mirrored anatomy described the same way. Nothing downstream cares: an +// even-odd fill has no winding, and ring SIMPLICITY is what is asserted. +export const EYE_R_RING = [ + 33, 246, 161, 160, 159, 158, 157, 173, + 133, 155, 154, 153, 145, 144, 163, 7, +]; +export const EYE_L_RING = [ + 263, 466, 388, 387, 386, 385, 384, 398, + 362, 382, 381, 380, 374, 373, 390, 249, +]; + +// Outer, inner corner per eye. All four are also in RIGID, and that is the +// point: the eye's reference frame is built only from landmarks that do not +// move under performance, so a blink cannot be mistaken for a change of gaze. +export const EYE_R_CORNERS = [33, 133]; +export const EYE_L_CORNERS = [263, 362]; + +// Upper and lower lid centres. Their separation over the corner distance is the +// openness signal that decides whether the eye is shut - the same shape of +// measurement as APERTURE is for the mouth, but normalised, so one threshold +// carries across takes and faces. +export const EYE_R_LIDS = [159, 145]; +export const EYE_L_LIDS = [386, 374]; + +// The two iris blocks the refined mesh appends: centre first, then four ring +// points. WHICH BLOCK BELONGS TO WHICH EYE IS NOT DECLARED HERE - MediaPipe's +// own "left"/"right" is viewer-relative in some docs and subject-relative in +// others, and a swap looks almost right, so it would survive an eyeball and +// then read as a permanently wall-eyed character. pipeline.js resolves it from +// the geometry instead. +export const IRIS_A = [468, 469, 470, 471, 472]; +export const IRIS_B = [473, 474, 475, 476, 477]; diff --git a/js/mathutil.js b/js/mathutil.js index 310eb8c..b49c51c 100644 --- a/js/mathutil.js +++ b/js/mathutil.js @@ -108,3 +108,27 @@ export function smoothTransforms(tfs, radius) { theta: Math.atan2(sn[i], c[i]), s: s[i], tx: tx[i], ty: ty[i], })); } + +// Push a ring outward from its centroid by a FIXED distance, not by a scale +// factor. +// +// Scaling collapses with the shape: a shut eyelid scaled by 1.1 is still a shut +// eyelid, so the lash line - the only thing left to draw when the eye is closed +// - would vanish exactly on the frames where it is the whole drawing. A fixed +// radial offset gives a band of roughly constant thickness that survives the +// ring going degenerate, and it keeps a star-shaped ring simple, which +// docs/design.md requires of every cut part. +export function offsetRing(pts, d) { + if (!d) return pts; + let cx = 0, cy = 0; + for (const p of pts) { cx += p.x; cy += p.y; } + cx /= pts.length; cy /= pts.length; + return pts.map((p) => { + const dx = p.x - cx, dy = p.y - cy; + const m = Math.hypot(dx, dy); + // A vertex sitting exactly on the centroid has no outward direction. Leave + // it where it is rather than emitting NaN and poisoning the whole ring. + return m < 1e-9 ? { x: p.x, y: p.y } + : { x: p.x + (dx / m) * d, y: p.y + (dy / m) * d }; + }); +} diff --git a/js/pipeline.js b/js/pipeline.js index 2e1155c..e004e72 100644 --- a/js/pipeline.js +++ b/js/pipeline.js @@ -2,7 +2,9 @@ // All policy lives here, never in the renderer. See docs/design.md, // "The take is the contract". -import { RIGID, LIPS_OUTER, LIPS_INNER, APERTURE, FACE_OVAL, EYE_INNER, subsampleSlots } from './landmarks.js'; +import { RIGID, LIPS_OUTER, LIPS_INNER, APERTURE, FACE_OVAL, EYE_INNER, + EYE_R_RING, EYE_L_RING, EYE_R_CORNERS, EYE_L_CORNERS, + EYE_R_LIDS, EYE_L_LIDS, IRIS_A, IRIS_B, subsampleSlots } from './landmarks.js'; import { fitSimilarity, applySimAll, applySim, fitResidual, procrustesMean, smoothTransforms, movingAverage } from './mathutil.js'; // MediaPipe normalises x by image WIDTH and y by image HEIGHT, so its normalised @@ -25,6 +27,13 @@ export function stabilize(dense, smoothRadius, aspect = 1) { const raw = rigid.map((r) => fitSimilarity(r, ref)); const tfs = smoothTransforms(raw, smoothRadius); + // The refined mesh appends ten iris points to the 468 face points, but a + // plain mesh does not, and synthetic or hand-fed tracks need not. Checked + // rather than assumed: reading past the end would surface as NaN gaze deep + // downstream instead of as "this track carries no iris". + const hasIris = dense.every((f) => f && f.length > IRIS_B[IRIS_B.length - 1]); + const map = (table) => dense.map((f, i) => applySimAll(tfs[i], pick(f, table, aspect))); + return { ref, transforms: tfs, @@ -43,9 +52,221 @@ export function stabilize(dense, smoothRadius, aspect = 1) { const a = applySimAll(tfs[i], pick(f, APERTURE, aspect)); return Math.hypot(a[0].x - a[1].x, a[0].y - a[1].y); }), + // Eyes. Lid rings are a feature and get traced like the mouth; corners and + // lid centres are the measurement frame; the iris blocks are raw until + // pairIrises decides which is which. + lidR: map(EYE_R_RING), lidL: map(EYE_L_RING), + cornersR: map(EYE_R_CORNERS), cornersL: map(EYE_L_CORNERS), + lidsR: map(EYE_R_LIDS), lidsL: map(EYE_L_LIDS), + irisA: hasIris ? map(IRIS_A) : null, + irisB: hasIris ? map(IRIS_B) : null, }; } +/* ---------- eyes ---------- */ + +const mid = (a, b) => ({ x: (a.x + b.x) / 2, y: (a.y + b.y) / 2 }); +const dist = (a, b) => Math.hypot(a.x - b.x, a.y - b.y); + +// Which iris block belongs to which eye is RESOLVED FROM THE DATA, not declared +// in a table. +// +// The naming in MediaPipe's own material is viewer-relative in some places and +// subject-relative in others, and the two blocks are otherwise +// indistinguishable. Getting it backwards swaps the irises, which looks almost +// right - each eye still has a disc in roughly the right place - so it survives +// a casual eyeball and then reads as a subtly wall-eyed character for the rest +// of the project. Proximity to the eye's corner midpoint settles it in one +// comparison, is impossible to get wrong, and keeps working if the model is +// ever renumbered. +// +// Voted across every frame rather than read off frame zero: one bad detection +// should not decide the whole shot. +export function pairIrises(stab) { + if (!stab.irisA) return null; + let votes = 0; + for (let f = 0; f < stab.irisA.length; f++) { + const cR = mid(stab.cornersR[f][0], stab.cornersR[f][1]); + votes += dist(stab.irisA[f][0], cR) < dist(stab.irisB[f][0], cR) ? 1 : -1; + } + return votes > 0 ? { right: 'irisA', left: 'irisB' } + : { right: 'irisB', left: 'irisA' }; +} + +// Per-frame eye measurements, in units of eye width. Measurement only - every +// threshold and every stylisation is applied by the callers. +// +// Everything here stays in HEAD-LOCAL space, which is the same space the mouth +// lives in and the same space the registered photo underlay is drawn in. An +// earlier version pinned each eye into a fixed socket fitted to its corners' +// mean over the shot. That does remove the wobble, but it removes too much: the +// residual from out-of-plane rotation is real motion of the eye relative to the +// head, it is still there in the footage, and pinning it away leaves the drawn +// eyes hanging still over a photo whose eyes are moving. The eye has to track +// the face exactly as the mouth does. +// +// The wobble the socket was aimed at is dealt with the way docs/design.md deals +// with it everywhere else - the bounded contour average, the same knob and the +// same radius the mouth uses - and by placing the iris in the frame of the +// ALREADY-SMOOTHED lid ring, so the iris cannot jitter independently of the eye +// it sits in. See buildEyes in app.js. +export function eyeSignals(stab) { + const N = stab.transforms.length; + const pairing = pairIrises(stab); + const openR = [], openL = [], gazeRaw = [], gazeR = [], gazeL = []; + + for (let f = 0; f < N; f++) { + const cR = mid(stab.cornersR[f][0], stab.cornersR[f][1]); + const cL = mid(stab.cornersL[f][0], stab.cornersL[f][1]); + const wR = dist(stab.cornersR[f][0], stab.cornersR[f][1]); + const wL = dist(stab.cornersL[f][0], stab.cornersL[f][1]); + + // Openness is the lid gap over the CORNER distance. Normalising by the + // corners rather than by anything derived from the lids keeps the + // denominator rigid, so the ratio measures the lid and nothing else, and + // one threshold carries across takes, faces and framings. + openR.push(dist(stab.lidsR[f][0], stab.lidsR[f][1]) / wR); + openL.push(dist(stab.lidsL[f][0], stab.lidsL[f][1]) / wL); + + if (!pairing) { + gazeRaw.push({ x: 0, y: 0 }); gazeR.push({ x: 0, y: 0 }); gazeL.push({ x: 0, y: 0 }); + continue; + } + const iR = stab[pairing.right][f][0], iL = stab[pairing.left][f][0]; + // Gaze is the iris centre relative to the CORNER MIDPOINT, in eye widths - + // a pure offset WITHIN the eye, with the eye's own position divided out, so + // that quantising it quantises the glance and not the head motion carrying + // it. + // + // Measuring against the lid ring's centroid instead would track the lid: + // every blink pulls that centroid down and would fake a glance at the + // floor, on precisely the frames where the eye is most conspicuous. The + // corners are in RIGID, so this origin and this denominator are both immune + // to the performance they are measuring. + const gR = { x: (iR.x - cR.x) / wR, y: (iR.y - cR.y) / wR }; + const gL = { x: (iL.x - cL.x) / wL, y: (iL.y - cL.y) / wL }; + + // ONE gaze for both eyes, and deliberately so. At 320x200 an iris is a + // handful of pixels and its centre comes from five landmarks on an eye + // twenty pixels wide, so the difference between the two measurements is + // noise, not vergence - and independent per-eye noise reads as wall-eyed + // immediately, which is the most expensive artefact on a face. Openness + // stays per-eye, because a wink is real performance and should survive. + gazeRaw.push({ x: (gR.x + gL.x) / 2, y: (gR.y + gL.y) / 2 }); + // Kept separately purely as a diagnostic. The two eyes should agree; when + // they disagree in a sustained way rather than frame to frame, that is not + // noise but out-of-plane head rotation biasing the projected iris offset, + // and no 2D measurement can undo it. + gazeR.push(gR); gazeL.push(gL); + } + return { openR, openL, gazeRaw, gazeR, gazeL, hasIris: !!pairing }; +} + +// Where "not looking anywhere in particular" sits on THIS face. Everything the +// character does is measured as a departure from it, so getting it wrong does +// not bias the gaze slightly - it re-points the whole performance. +// +// `median` is the default and the safe one: the middle of the take, per axis. +// docs/design.md already gives this rule for the anchor fit - the reference is +// the MEAN configuration over the shot, not one frame - and gaze needs it for +// the same reason. The median rather than the mean because a couple of frames +// of hard glance should not drag the rest-point after them. +// +// `neutral` reads the origin off the take's neutral frame instead, which is +// only correct when there genuinely is a held neutral to read. That frame is +// chosen by MINIMUM MOUTH APERTURE, and a closed mouth says nothing whatever +// about where the eyes are pointed - so on footage with no deliberate neutral +// at the top it is an arbitrary frame, and whichever way the performer happened +// to glance on it becomes "straight ahead" for the entire shot. It is kept +// because it is right when the take was shot for this tool, and because being +// able to switch is how you find out that it was not. +export function gazeOrigin(gazeRaw, mode = 'median', neutral = 0, radius = 2) { + if (mode === 'neutral') { + let sx = 0, sy = 0, n = 0; + // A window, not a single frame: one frame of a five-landmark iris centre is + // worth about a pixel of noise, and that pixel would become a permanent + // squint in the output. + for (let f = neutral - radius; f <= neutral + radius; f++) { + const k = Math.min(gazeRaw.length - 1, Math.max(0, f)); + sx += gazeRaw[k].x; sy += gazeRaw[k].y; n++; + } + return { x: sx / n, y: sy / n }; + } + const mid1 = (vals) => { + const v = vals.slice().sort((a, b) => a - b); + return v.length % 2 ? v[(v.length - 1) / 2] + : (v[v.length / 2 - 1] + v[v.length / 2]) / 2; + }; + return { x: mid1(gazeRaw.map((g) => g.x)), y: mid1(gazeRaw.map((g) => g.y)) }; +} + +// Snap gaze onto a grid, then require a new cell to hold before it takes. +// +// This is the "Primitive - quantised" row of the part table in docs/design.md, +// and it is not a stylisation imposed on the truth: real eyes move in saccades, +// holding a fixation and then jumping. The smooth drift left in the measurement +// is tracker noise plus head-compensation error, so snapping to a grid and +// requiring a dwell removes the noise and recovers the saccade in the same +// operation - the rare case where the aesthetic rule and the physiology agree. +// +// The dwell is what stops a gaze parked on a cell boundary from chattering +// between two cells forever. It is meaningless without a grid, because +// continuous values never repeat, so step 0 short-circuits both. +export function quantizeGaze(gaze, step, dwell) { + if (!(step > 0)) return gaze.map((g) => ({ x: g.x, y: g.y })); + const q = gaze.map((g) => ({ + x: Math.round(g.x / step) * step, + y: Math.round(g.y / step) * step, + })); + if (dwell <= 0 || !q.length) return q; + + const out = []; + let live = q[0], pend = q[0], run = 0; + for (const g of q) { + if (g.x === pend.x && g.y === pend.y) run++; + else { pend = g; run = 1; } + if (run > dwell && (pend.x !== live.x || pend.y !== live.y)) live = pend; + out.push(live); + } + return out; +} + +// Resolve openness into a shut/open decision per frame. +// +// `dwell` is the same guard the teeth get: a lid hovering at the threshold must +// commit before the state changes, so it cannot flicker. +// +// `hold` is the one that is NOT like the teeth, and it is the whole reason +// blinks are worth special-casing. A blink is 100-150ms, which at 12fps is one +// frame and at 24fps is two or three - and a single frame of closed eye reads +// as a dropped frame, not as a blink. Animators draw a blink over two or three +// drawings for exactly that reason. So once the eye shuts it stays shut for +// `hold` frames, which turns an unreadable flicker into a beat. +// +// The hysteresis runs the other way from the teeth: shutting needs a clear +// signal, and once shut the eye is given the benefit of the doubt on reopening, +// because the lid landmarks are least reliable mid-blink. +export function resolveBlink(open, { cut, dwell, hold }) { + const N = open.length; + const shut = new Array(N).fill(false); + let live = false; // current state + let run = 0; // frames the opposing reading has persisted + let held = 0; // frames spent in the current state + for (let f = 0; f < N; f++) { + const reading = live ? open[f] < cut * 1.35 : open[f] < cut; + if (reading === live) run = 0; + else { + run++; + // Leaving a blink additionally requires the blink to have been on screen + // long enough to be legible; entering one never waits. + if (run > dwell && (!live || held >= hold)) { live = reading; held = 0; run = 0; } + } + held++; + shut[f] = live; + } + return shut; +} + // Stage 4: fixed-index subsample of a stabilised ring, then map from normalised // face space into character raster space. export function toRasterRing(stabRing, ringTable, n, xform) { @@ -178,6 +399,24 @@ export function heldFrame(kept, f) { return hit; } +// Hold every output frame back onto an exposure grid: 1 = on 1s, 2 = on 2s, and +// so on. Frame 5 at exposure 2 reads the pose from frame 4. +// +// This is where "aesthetic sparseness" belongs. docs/design.md used to put it at +// the extraction rate - pick 12fps and the timing is already chosen - but that +// makes the timing a property of a directory of PNGs, so auditioning 12 against +// 24 means re-ripping the clip and re-running detection over all of it. Rip +// dense once and quantise here instead: the dense track stays at the camera's +// rate, the decision stays reversible, and the audio clock is untouched, so +// sync cannot drift while you try timings. +// +// Floor, never round. Rounding would let an output frame read a pose from the +// FUTURE, which is a lead - a separate control, applied after this one, for a +// separate reason. +export function exposeIndex(f, exposure) { + return exposure > 1 ? Math.floor(f / exposure) * exposure : f; +} + // Shift a performance track against the clock, clamped at the ends. // // Pure and exported so the shift can actually be asserted: "the slider feels diff --git a/js/raster.js b/js/raster.js index 477e63f..987fbe8 100644 --- a/js/raster.js +++ b/js/raster.js @@ -46,14 +46,47 @@ export class IndexedRaster { } } - fillDisc(cx, cy, r, index) { + // `over` is an optional stencil: when given, only pixels that currently hold + // that index are written. The indexed buffer is its own clip mask, which is + // how Animator Pro would do it - and it is what keeps the iris inside the + // eye. A disc clipped by the sclera cannot spill past the lid at any gaze or + // any radius, including mid-blink when the opening is a two-pixel sliver, so + // the lid crops the iris for free instead of the gaze range needing a + // clamp that would flatten the performance at the extremes. + fillDisc(cx, cy, r, index, over = null) { const rr = r * r; const y0 = Math.max(0, Math.floor(cy - r)), y1 = Math.min(this.h - 1, Math.ceil(cy + r)); const x0 = Math.max(0, Math.floor(cx - r)), x1 = Math.min(this.w - 1, Math.ceil(cx + r)); for (let y = y0; y <= y1; y++) { for (let x = x0; x <= x1; x++) { const dx = x + 0.5 - cx, dy = y + 0.5 - cy; - if (dx * dx + dy * dy <= rr) this.buf[y * this.w + x] = index; + if (dx * dx + dy * dy > rr) continue; + const o = y * this.w + x; + if (over === null || this.buf[o] === over) this.buf[o] = index; + } + } + } + + // An exactly size x size block of pixels, snapped to the pixel grid, with the + // same optional stencil as fillDisc. + // + // The pupil is a SQUARE because at 320x200 it is three pixels across, and a + // circle of radius 1.5 is not a circle - it is a plus sign with the corners + // gnawed off, and it changes shape as it moves. A square that size is a + // deliberate mark that stays the same mark wherever it lands, which is the + // whole argument for flat shapes at this resolution. + // + // The top-left is rounded rather than the centre, so the block is size x size + // on every frame. Round the extents instead and a fractional centre gives you + // three pixels on one frame and four on the next, which reads as the pupil + // breathing. + fillRect(cx, cy, size, index, over = null) { + if (size < 1) return; + const x0 = Math.round(cx - size / 2), y0 = Math.round(cy - size / 2); + for (let y = Math.max(0, y0); y < Math.min(this.h, y0 + size); y++) { + for (let x = Math.max(0, x0); x < Math.min(this.w, x0 + size); x++) { + const o = y * this.w + x; + if (over === null || this.buf[o] === over) this.buf[o] = index; } } } diff --git a/js/selftest.js b/js/selftest.js index 026c476..ed58aa4 100644 --- a/js/selftest.js +++ b/js/selftest.js @@ -7,9 +7,12 @@ // as blocks meeting at corners. It is invisible at some vertex counts and obvious // at others, so it needs an assertion rather than an eyeball. -import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, subsampleSlots, subsampleRing } from './landmarks.js'; -import { fitSimilarity, applySim, procrustesMean, smoothTransforms } from './mathutil.js'; -import { stabilize, toRasterRing, selectKeys, activeKey, shiftIndex } from './pipeline.js'; +import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, subsampleSlots, subsampleRing, + EYE_R_RING, EYE_L_RING, EYE_R_CORNERS, EYE_L_CORNERS, + EYE_R_LIDS, EYE_L_LIDS } from './landmarks.js'; +import { fitSimilarity, applySim, procrustesMean, smoothTransforms, offsetRing } from './mathutil.js'; +import { stabilize, toRasterRing, smoothContours, selectKeys, activeKey, shiftIndex, + exposeIndex, eyeSignals, pairIrises, gazeOrigin, quantizeGaze, resolveBlink } from './pipeline.js'; import { IndexedRaster, hexToRgb } from './raster.js'; import { writeTake } from './take.js'; import { otsuForTest, scaleRing } from './interior.js'; @@ -206,6 +209,25 @@ export function run() { ok('activeKey holds between keys', activeKey(sel.keys, sel.keys[1].f - 1).f === sel.keys[0].f); + // exposure: rip dense, choose the timing here. On 2s every odd frame must + // reuse the even frame's pose, and the grid must never read from the future - + // that direction is the lead, which is a different control for a reason. + ok('exposure 1 is identity', [0, 1, 7, 71].every((f) => exposeIndex(f, 1) === f)); + ok('on 2s holds each pose for two frames', + [0, 1, 2, 3, 4, 5].map((f) => exposeIndex(f, 2)).join(',') === '0,0,2,2,4,4'); + ok('on 3s holds each pose for three frames', + [0, 1, 2, 3, 4, 5, 6].map((f) => exposeIndex(f, 3)).join(',') === '0,0,0,3,3,3,6'); + ok('exposure never reads a pose from the future', + [0, 1, 2, 3, 4, 5, 6, 7].every((f) => exposeIndex(f, 3) <= f)); + { + // Exposure then lead, in that order: the picture must change on the grid + // beats and carry a pose shifted by whole frames of the original track. + const N = 72, at = (f) => shiftIndex(exposeIndex(f, 2), 1, N); + ok('exposure and lead compose without moving the beats', + at(0) === 1 && at(1) === 1 && at(2) === 3 && at(3) === 3, + [0, 1, 2, 3].map(at).join(',')); + } + // mouth lead: a shift that "feels like it does nothing" is indistinguishable // from one that does nothing, so assert the arithmetic directly. ok('lead 0 is identity', [0, 5, 71].every((f) => shiftIndex(f, 0, 72) === f)); @@ -288,6 +310,241 @@ export function run() { (tt.mBright - tt.mDark) / 255 > 0.4, `sep ${((tt.mBright - tt.mDark) / 255).toFixed(4)}`); } + /* ---- eyes ---- */ + + ok('eye rings have 16 distinct ids each', + new Set(EYE_R_RING).size === 16 && new Set(EYE_L_RING).size === 16); + ok('the two eye rings share no landmark', + !EYE_R_RING.some((i) => EYE_L_RING.includes(i))); + + // The cardinal contract, asserted rather than trusted: on a 16-slot ring the + // quarter slots must be the four anatomical cardinals, which is what makes + // every even vertex budget land on real landmarks instead of between them. + ok('eye ring slot 0/4/8/12 are outer, upper, inner, lower', + EYE_R_RING[0] === EYE_R_CORNERS[0] && EYE_R_RING[8] === EYE_R_CORNERS[1] && + EYE_R_RING[4] === EYE_R_LIDS[0] && EYE_R_RING[12] === EYE_R_LIDS[1] && + EYE_L_RING[0] === EYE_L_CORNERS[0] && EYE_L_RING[8] === EYE_L_CORNERS[1] && + EYE_L_RING[4] === EYE_L_LIDS[0] && EYE_L_RING[12] === EYE_L_LIDS[1]); + + // The gaze origin and denominator are built from the eye corners, so if a + // corner were not rigid a blink could move it and fake a glance. + ok('every eye corner is a rigid landmark', + [...EYE_R_CORNERS, ...EYE_L_CORNERS].every((i) => RIGID.includes(i))); + + // Same simplicity requirement as the lips, and for the same reason: a cut + // part with a self-intersecting ring renders as blocks meeting at corners. + // Checked on blink frames too, where the ring is nearly degenerate. + for (const [label, table] of [['right', EYE_R_RING], ['left', EYE_L_RING]]) { + let worst = null; + for (let n = 4; n <= 12 && !worst; n += 2) { + const slots = subsampleSlots(table.length, n); + for (let f = 0; f < dense.length; f++) { + const hits = ringSelfIntersections(slots.map((sl) => dense[f][table[sl]])); + if (hits.length) { worst = `verts=${n} frame=${f} edges ${JSON.stringify(hits[0])}`; break; } + } + } + ok(`${label} eye ring is simple at every vertex budget`, !worst, worst || ''); + } + + // offsetRing must grow by a FIXED amount and survive a degenerate ring - the + // shut eyelid is exactly the degenerate case, and it is the frame where the + // lash line is the entire drawing. + { + const sq = [{ x: -1, y: 0 }, { x: 0, y: -1 }, { x: 1, y: 0 }, { x: 0, y: 1 }]; + const g = offsetRing(sq, 2); + ok('offsetRing pushes every vertex out by exactly d', + g.every((p, i) => Math.abs(Math.hypot(p.x, p.y) - (Math.hypot(sq[i].x, sq[i].y) + 2)) < 1e-9)); + ok('offsetRing(0) is identity', offsetRing(sq, 0) === sq); + // A shut lid: a flat sliver. The offset must still open it into a band. + const shutLid = [{ x: -10, y: 0 }, { x: 0, y: -0.02 }, { x: 10, y: 0 }, { x: 0, y: 0.02 }]; + const band = offsetRing(shutLid, 1.5); + const h = Math.max(...band.map((p) => p.y)) - Math.min(...band.map((p) => p.y)); + ok('offsetRing gives a shut lid a visible lash band', h > 2.9, `height ${h.toFixed(3)}`); + ok('offsetRing keeps the shut lid simple', ringSelfIntersections(band).length === 0); + } + + { + const stE = stabilize(dense, 2); + const sig = eyeSignals(stE); + ok('synthetic track carries iris landmarks', sig.hasIris); + + // THE load-bearing eye assertion. The pairing is resolved from geometry + // rather than declared, so the test feeds a track built the OTHER way round + // and demands the resolver follow the data. A resolver only ever checked + // against the convention it was written for is checking nothing. + const pairA = pairIrises(stE); + const pairB = pairIrises(stabilize(synthDense(72, { swapIris: true }), 2)); + ok('iris pairing is resolved from the data, not assumed', + pairA.right === 'irisA' && pairB.right === 'irisB', + `normal ${pairA.right}, swapped ${pairB.right}`); + + // The eye must TRACK the face, not sit in a fixed socket. An earlier + // version pinned each eye to its corners' mean over the shot, which does + // kill the wobble but leaves the drawn eyes hanging still over a registered + // photo whose eyes are moving. Head-local is the same space the mouth and + // the underlay live in, so the eye moves with the head exactly as they do. + { + const spread = (arr, sel) => { + const v = arr.map(sel); + return Math.max(...v) - Math.min(...v); + }; + const w = Math.hypot(stE.cornersR[0][0].x - stE.cornersR[0][1].x, + stE.cornersR[0][0].y - stE.cornersR[0][1].y); + const moves = Math.max(spread(stE.lidR, (r) => r[0].x), spread(stE.lidR, (r) => r[0].y)); + ok('the eye stays in head-local space and tracks the face', moves / w > 0.02, + `corner travels ${(moves / w * 100).toFixed(1)}% of an eye width`); + + // Subsampling a 16-slot ring to any even budget must keep the two corners + // at output indices 0 and n/2. That is what lets the socket be read back + // off the drawn polygon instead of measured separately, which is what + // stops the iris drifting relative to the eye it sits in. + let bad = null; + for (let n = 4; n <= 12; n += 2) { + const sl = subsampleSlots(16, n); + if (sl[0] !== 0 || sl[n / 2] !== 8) bad = `n=${n} -> ${sl.join(',')}`; + } + ok('the drawn lid ring carries its own corners at 0 and n/2', !bad, bad || ''); + + // The contour average is what removes the jitter, and it is the mouth's + // knob doing the mouth's job - no second mechanism for the eyes. + const ring = (rad) => smoothContours( + stE.lidR.map((r) => toRasterRing(r, EYE_R_RING, 8, (p) => ({ x: p.x * 600, y: p.y * 600 }))), rad); + const jitter = (rings) => { + let acc = 0; + for (let f = 1; f < rings.length; f++) { + const a = rings[f], b = rings[f - 1]; + acc += Math.hypot((a[0].x + a[4].x) / 2 - (b[0].x + b[4].x) / 2, + (a[0].y + a[4].y) / 2 - (b[0].y + b[4].y) / 2); + } + return acc / (rings.length - 1); + }; + ok('contour averaging steadies the eye without pinning it', + jitter(ring(1)) < jitter(ring(0)) * 0.8, + `${jitter(ring(0)).toFixed(3)} -> ${jitter(ring(1)).toFixed(3)} px/frame`); + } + + // Blink: synth shuts the lids for exactly one frame every 19. + const lo = Math.min(...sig.openR), hi = Math.max(...sig.openR); + ok('openness collapses on a blink and not otherwise', lo < hi * 0.2, + `${lo.toFixed(3)} .. ${hi.toFixed(3)}`); + + const shut = resolveBlink(sig.openR, { cut: hi * 0.3, dwell: 0, hold: 3 }); + const runs = []; + for (let f = 0; f < shut.length; f++) if (shut[f] && !shut[f - 1]) runs.push(f); + const lens = runs.map((a) => { let n = 0; while (shut[a + n]) n++; return n; }); + ok('blinks are found', runs.length >= 3, `${runs.length} runs at ${runs.join(',')}`); + // The knob that is not like the teeth: a one-frame blink reads as a dropped + // frame, so `hold` must stretch it into something legible. + ok('a one-frame blink is held to the minimum length', + lens.every((n) => n >= 3), `run lengths ${lens.join(',')}`); + ok('a shorter hold leaves the blink shorter', + resolveBlink(sig.openR, { cut: hi * 0.3, dwell: 0, hold: 1 }).filter(Boolean).length < + shut.filter(Boolean).length); + + // Gaze, against ground truth: synth commands +0.16 eye widths at f12 and + // -0.16 at f23, holding each for eleven frames. + const org = gazeOrigin(sig.gazeRaw, 'neutral', 0); + const gx = (f) => (sig.gazeRaw[f].x - org.x); + ok('gaze recovers the commanded direction', + gx(12) > 0.12 && gx(12) < 0.20 && gx(23) < -0.12 && gx(23) > -0.20, + `f12 ${gx(12).toFixed(3)}, f23 ${gx(23).toFixed(3)}`); + + // Measuring gaze against the lid centroid instead of the corner midpoint + // would drag the iris down on every blink and fake a glance at the floor, + // on exactly the frames where the eye is most conspicuous. + const gy = (f) => (sig.gazeRaw[f].y - org.y); + ok('a blink does not fake a change of gaze', + Math.abs(gy(19) - gy(18)) < 0.02, `f18 ${gy(18).toFixed(4)} -> f19 ${gy(19).toFixed(4)}`); + + // Quantisation is what turns drift into saccades: four commanded + // fixations must come back as a handful of cells, not one per frame. + // The origin re-points the whole performance, so a wrong one does not bias + // the gaze slightly - it makes the character look the other way. The median + // must sit inside the range it summarises; the neutral-frame origin need + // not, which is exactly the failure mode it has on footage with no + // deliberate neutral at the top. + { + const med = gazeOrigin(sig.gazeRaw, 'median'); + const xs = sig.gazeRaw.map((g) => g.x); + ok('the median origin lies inside the take\'s own gaze range', + med.x > Math.min(...xs) && med.x < Math.max(...xs), + `${med.x.toFixed(3)} in ${Math.min(...xs).toFixed(3)}..${Math.max(...xs).toFixed(3)}`); + // Synth looks left as much as right, so the rest point is near zero. + ok('the median origin finds the rest point, not a glance', + Math.abs(med.x) < 0.08, `median x ${med.x.toFixed(3)}`); + ok('the two origins actually differ, so the toggle is a real A/B', + Math.abs(med.x - gazeOrigin(sig.gazeRaw, 'neutral', 12).x) > 0.02); + } + + const px = sig.gazeRaw.map((g) => ({ x: (g.x - org.x) * 30, y: (g.y - org.y) * 30 })); + const cells = (a) => new Set(a.map((g) => `${g.x},${g.y}`)).size; + ok('quantisation collapses drift into a few fixations', + cells(quantizeGaze(px, 2, 2)) <= 6 && cells(px) > 40, + `${cells(px)} raw -> ${cells(quantizeGaze(px, 2, 2))} cells`); + ok('gaze step 0 leaves the track untouched', + quantizeGaze(px, 0, 2).every((g, i) => g.x === px[i].x && g.y === px[i].y)); + ok('quantised values land on the grid', + quantizeGaze(px, 2, 0).every((g) => Math.abs(g.x % 2) < 1e-9 && Math.abs(g.y % 2) < 1e-9)); + // A one-frame excursion is noise; the dwell must swallow it. + { + const spike = [{ x: 0, y: 0 }, { x: 0, y: 0 }, { x: 4, y: 0 }, { x: 0, y: 0 }, { x: 0, y: 0 }]; + ok('the dwell suppresses a one-frame gaze spike', + quantizeGaze(spike, 2, 1).every((g) => g.x === 0)); + ok('a sustained move still gets through', + quantizeGaze([...spike, { x: 4, y: 0 }, { x: 4, y: 0 }, { x: 4, y: 0 }], 2, 1).pop().x === 4); + } + } + + // The iris is stencilled by the sclera and the pupil by the iris, which is + // what keeps both inside the lid at any gaze without clamping the gaze itself. + { + const rr = new IndexedRaster(40, 40); + rr.clear(0); + rr.fillPoly([{ x: 10, y: 10 }, { x: 30, y: 10 }, { x: 30, y: 20 }, { x: 10, y: 20 }], 1); + rr.fillDisc(28, 15, 9, 2, 1); // a disc reaching well past the "lid" + let spill = 0, inside = 0; + for (let y = 0; y < 40; y++) for (let x = 0; x < 40; x++) { + const v = rr.buf[y * 40 + x]; + if (v !== 2) continue; + if (x >= 10 && x < 30 && y >= 10 && y < 20) inside++; else spill++; + } + ok('a stencilled disc cannot spill past its clip', spill === 0 && inside > 20, + `${inside} in, ${spill} out`); + rr.fillDisc(5, 35, 3, 3); // no stencil: writes freely + ok('an unstencilled disc still writes anywhere', rr.buf.includes(3)); + + // The stencil chain: pupil over iris over sclera. A pupil placed where the + // iris has already been cropped must be cropped the same way. + rr.fillRect(28, 15, 5, 4, 2); + let pSpill = 0; + for (let y = 0; y < 40; y++) for (let x = 0; x < 40; x++) { + if (rr.buf[y * 40 + x] === 4 && !(x >= 10 && x < 30 && y >= 10 && y < 20)) pSpill++; + } + ok('the pupil inherits the iris clip transitively', pSpill === 0); + } + + // A square pupil is only worth having if it is the SAME square every frame: + // exactly its nominal size at any centre, or it breathes as the gaze moves. + { + const sizes = []; + for (const [cx, cy] of [[20, 20], [20.5, 20.5], [20.49, 19.51], [21, 20]]) { + const rr = new IndexedRaster(40, 40); + rr.clear(0); + rr.fillRect(cx, cy, 3, 1); + let n = 0, minX = 99, maxX = -1, minY = 99, maxY = -1; + for (let y = 0; y < 40; y++) for (let x = 0; x < 40; x++) { + if (rr.buf[y * 40 + x] !== 1) continue; + n++; minX = Math.min(minX, x); maxX = Math.max(maxX, x); + minY = Math.min(minY, y); maxY = Math.max(maxY, y); + } + sizes.push(`${maxX - minX + 1}x${maxY - minY + 1}:${n}`); + } + ok('a 3px pupil is 3x3 at every centre', sizes.every((v) => v === '3x3:9'), sizes.join(' ')); + const rr = new IndexedRaster(40, 40); + rr.clear(0); rr.fillRect(20, 20, 0, 1); + ok('pupil size 0 draws nothing', !rr.buf.includes(1)); + } + // take writer round-trip const take = { name: 'test', frames: 72, width: 320, height: 200, exposure: 2, @@ -297,6 +554,12 @@ export function run() { { name: 'head', kind: 'plate', z: 0, interp: 'hold', keys: [{ f: 0, plate: 0 }] }, { name: 'mouth', kind: 'poly', z: 30, color: 'skin', interp: 'hold', keys: sel.keys.map((k) => ({ f: k.f, src: k.src, pts: shapes[k.src] })) }, + { name: 'iris_r', kind: 'disc', z: 22, color: 'iris', interp: 'hold', + parent: 'eye_r_in', clip: 'eye_r_in', + keys: [{ f: 0, src: 0, c: { x: 120.4, y: 88.7 }, r: 5.5 }, { f: 2, hidden: true }] }, + { name: 'pupil_r', kind: 'rect', z: 23, color: 'pupil', interp: 'hold', + parent: 'iris_r', clip: 'iris_r', + keys: [{ f: 0, src: 0, c: { x: 120, y: 89 }, size: 3 }] }, ], }; const text = writeTake(take); @@ -309,6 +572,13 @@ export function run() { }), `${keyLines.length} key lines`); ok('take declares a plate and a part table', /^plate\s+0/m.test(text) && /^part\s+mouth/m.test(text)); + ok('a disc part declares its clip', /^part\s+iris_r.*clip=eye_r_in/m.test(text)); + ok('a disc key is three integers', /^key\s+iris_r\s+f=0\s+src=0\s+disc=120,89,6$/m.test(text), + (text.split('\n').find((l) => l.startsWith('key iris_r')) || '').trim()); + ok('a hidden disc key emits hidden', /^key\s+iris_r\s+f=2\s+hidden$/m.test(text)); + ok('a pupil key is a square, not a tessellated polygon', + /^key\s+pupil_r\s+f=0\s+src=0\s+rect=120,89,3$/m.test(text) && + /^part\s+pupil_r.*clip=iris_r/m.test(text)); ok('coordinates are integers', !/-?\d+\.\d/.test(text.split('\n').filter((l) => l.startsWith('key')).join(''))); return results; diff --git a/js/synth.js b/js/synth.js index bf736be..6de5606 100644 --- a/js/synth.js +++ b/js/synth.js @@ -5,11 +5,16 @@ // verified without a video file. A synthetic face is also the only way to test // stabilisation against a KNOWN head motion, since real footage gives no ground // truth to compare against. -import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, EYE_INNER } from './landmarks.js'; +import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, + EYE_R_RING, EYE_L_RING, IRIS_A, IRIS_B } from './landmarks.js'; const NUM = 478; -export function synthDense(nFrames = 72) { +// `swapIris` places the two iris blocks on the opposite eyes. It exists so the +// pairing resolver can be tested against a track it actually disagrees with: +// a resolver checked only against the convention it was written for is checking +// nothing at all. +export function synthDense(nFrames = 72, { swapIris = false } = {}) { const frames = []; for (let t = 0; t < nFrames; t++) { const pts = new Array(NUM); @@ -35,11 +40,49 @@ export function synthDense(nFrames = 72) { const openAmt = target[beat]; const wide = 0.10 + (beat === 1 ? 0.012 : beat === 3 ? -0.008 : 0); - place(RIGID[0], -0.075, -0.045); place(RIGID[1], -0.028, -0.043); - place(RIGID[2], 0.028, -0.043); place(RIGID[3], 0.075, -0.045); place(RIGID[4], 0.000, -0.050); place(RIGID[5], 0.000, -0.020); place(RIGID[6], 0.000, 0.012); - place(EYE_INNER[0], -0.028, -0.043); place(EYE_INNER[1], 0.028, -0.043); + + // Eyes. The corners (RIGID[0..3]) are placed BY the lid rings rather than + // separately, because they are slots 0 and 8 of those rings: writing them + // twice is how the mouth grew a bowtie, and a corner that disagrees with + // its own ring would make the eye self-intersect at some vertex budgets + // and not others. + // + // A blink is ONE frame, which is the honest hard case: at 12fps that is + // what a real blink costs, and it is exactly the length that reads as a + // dropped frame rather than as a blink unless `hold` extends it. + const blink = t > 5 && t % 19 === 0; + const openness = blink ? 0.05 : 1; + + // Gaze holds and then jumps, the way gaze actually behaves, with a little + // jitter on top so quantisation has noise to remove and the dwell has + // something to suppress. + const LOOK = [[0, 0], [0.16, 0.0], [-0.16, 0.05], [0.0, -0.09]]; + const [gx, gy] = LOOK[Math.floor(t / 11) % LOOK.length]; + const jit = () => (Math.random() - 0.5) * 0.012; + + // Half the corner separation, and the lid half-height at full open. + const EYE_RX = 0.0235, EYE_RY = 0.011, EYE_Y = -0.044; + const eye = (ring, cx, dir, iris) => { + const n = ring.length; + for (let k = 0; k < n; k++) { + // dir flips the traversal so each ring runs the direction its real + // table does: slot 0 outer corner, 4 upper lid, 8 inner, 12 lower. + const a = dir > 0 ? Math.PI + (k / n) * Math.PI * 2 : -(k / n) * Math.PI * 2; + place(ring[k], cx + EYE_RX * Math.cos(a), + EYE_Y + EYE_RY * openness * Math.sin(a)); + } + // Iris: centre first, then four ring points, as the refined mesh emits. + const ix = cx + (gx + jit()) * EYE_RX * 2, iy = EYE_Y + (gy + jit()) * EYE_RX * 2; + place(iris[0], ix, iy); + for (let k = 1; k < iris.length; k++) { + const a = ((k - 1) / (iris.length - 1)) * Math.PI * 2; + place(iris[k], ix + 0.008 * Math.cos(a), iy + 0.008 * Math.sin(a)); + } + }; + eye(EYE_R_RING, -0.0515, 1, swapIris ? IRIS_B : IRIS_A); + eye(EYE_L_RING, 0.0515, -1, swapIris ? IRIS_A : IRIS_B); // Lip rings as ellipse arcs, traversed so ring ORDER matches the tables: // slot 0 = right corner, 5 = top centre, 10 = left corner, 15 = bottom diff --git a/js/take.js b/js/take.js index 8a87b09..97929f3 100644 --- a/js/take.js +++ b/js/take.js @@ -15,7 +15,12 @@ export function writeTake(take) { // v1 emits a single frozen plate derived from the face oval. A real project // replaces this with hand-drawn angles referenced by cel frame; the record // shape is the same either way. - L.push(`plate 0 kind=poly slot_mouth=${r(take.slot.x)},${r(take.slot.y)} scale=1.00 rot=0 squash=1.00`); + L.push(`plate 0 kind=poly slot_mouth=${r(take.slot.x)},${r(take.slot.y)}` + + (take.eyeSlots + ? ` slot_eye_r=${r(take.eyeSlots.r.cx)},${r(take.eyeSlots.r.cy)},${r(take.eyeSlots.r.w)}` + + ` slot_eye_l=${r(take.eyeSlots.l.cx)},${r(take.eyeSlots.l.cy)},${r(take.eyeSlots.l.w)}` + : '') + + ` scale=1.00 rot=0 squash=1.00`); L.push(''); for (const part of take.parts) { const bits = [`part ${part.name.padEnd(9)} kind=${part.kind} z=${part.z}`]; @@ -23,6 +28,10 @@ export function writeTake(take) { if (part.kind === 'poly') bits.push('closed=1 fill=1'); bits.push(`interp=${part.interp}`); if (part.parent) bits.push(`parent=${part.parent}`); + // `clip` names a part this one is stencilled by, not merely drawn after. + // The iris needs it: at an extreme gaze the disc reaches past the lid, and + // ordering alone would put it on the cheek. + if (part.clip) bits.push(`clip=${part.clip}`); L.push(bits.join(' ')); } L.push(''); @@ -30,6 +39,20 @@ export function writeTake(take) { for (const k of part.keys) { if (k.hidden) { L.push(`key ${part.name.padEnd(9)} f=${k.f} hidden`); continue; } if (part.kind === 'plate') { L.push(`key ${part.name.padEnd(9)} f=${k.f} plate=${k.plate}`); continue; } + // A disc is three numbers, so it gets its own key shape rather than being + // pre-tessellated into a polygon here: the renderer draws a real circle + // with hard edges, and a five-pixel iris approximated by a polygon would + // lose a pixel off its silhouette on some frames and not others. + // A square, in whole pixels, at an integer centre. Same argument as the + // disc: the renderer is told the shape, not a polygon approximating it. + if (part.kind === 'rect') { + L.push(`key ${part.name.padEnd(9)} f=${k.f} src=${k.src} rect=${r(k.c.x)},${r(k.c.y)},${k.size}`); + continue; + } + if (part.kind === 'disc') { + L.push(`key ${part.name.padEnd(9)} f=${k.f} src=${k.src} disc=${r(k.c.x)},${r(k.c.y)},${r(k.r)}`); + continue; + } const pts = k.pts.map((p) => `${r(p.x)},${r(p.y)}`).join(' '); // f is authoritative (what renders); src is the pre-snap extreme frame, // kept only as a tuning signal - a key dragged far means the minimum-hold