2026-09-24 14:38:07 -04:00
|
|
|
// Assertions over the stages below detection. Runs in the browser so the exact
|
|
|
|
|
// module graph the tool uses is what gets tested.
|
|
|
|
|
//
|
|
|
|
|
// The ring-simplicity check exists because "fixed topology" is load-bearing in
|
Become arthur: a standalone suite, not an Animator Pro front-end
The test renderer turned out to be the product. Everything that decides how the
work looks - stabilisation, reduction, timing, frame removal, palette - already
happens here, and the flat indexed output already reads the way it should.
The reason to leave is in the original design's own rule: never make a timing
decision that requires a full render to evaluate. Honouring that moved every
judgement out of Animator Pro, which left the host doing nothing but writing a
file, in exchange for modal UI, minutes-long renders, one-level undo, FLX delta
invariants, a single tween state and a cel singleton.
What does NOT change is the constraint. 320x200, indexed palette, flat fills,
no antialiasing - inherited, but load-bearing rather than accidental. The
rasteriser writes palette indices and expands to RGBA only at the end precisely
so nothing can soften an edge. Modern conveniences belong in the workflow.
Adds docs/design.md: the principles, carried over without the Poco/FLX/cel
machinery, plus architecture and an honest list of what is missing - the
largest gap being that plates still have nowhere to be drawn.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:47:41 -04:00
|
|
|
// docs/design.md: because hold parts CUT between poses rather than
|
2026-09-24 14:38:07 -04:00
|
|
|
// interpolating, a ring whose vertex order is wrong self-intersects and renders
|
|
|
|
|
// as blocks meeting at corners. It is invisible at some vertex counts and obvious
|
|
|
|
|
// at others, so it needs an assertion rather than an eyeball.
|
|
|
|
|
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, subsampleSlots, subsampleRing,
|
|
|
|
|
EYE_R_RING, EYE_L_RING, EYE_R_CORNERS, EYE_L_CORNERS,
|
Brows: traced ring, quantised raise
A brow at 320x200 is fourteen pixels wide and three tall. Its shape carries
almost nothing at that size; its height above the eye carries the expression,
and a brow raise is the most legible beat on a face. So the ring is traced and
the height is quantised - the split the eyes already got, where the lid is a
traced feature and the iris a quantised primitive.
The decomposition is the point. The traced ring already contains the real
height, so adding a quantised raise on top would move the brow twice. The
height is measured OUT of the ring, quantised, and put back, so the shape that
renders is his at a height that snaps between a few levels and holds.
Measured at both ends rather than as one number, because raise and tilt are
different expressions out of one mechanism: both ends up is surprise, inner up
alone is worry, inner down is anger. They share a dwell - the gaze quantiser,
renamed quantizeSnap now that it has two callers - so the brow hits its pose in
one frame instead of crawling into it with one end arriving before the other.
Measured against the eye's corner midpoint, never its lid. Same trap the gaze
origin has and worth avoiding twice: brows and lids move together constantly,
so a brow that jumped on every blink would read as a tic. Rest pose from the
take median rather than the neutral frame, for the reason gaze learned the hard
way - that frame is picked by minimum mouth aperture and says nothing about the
brows.
Two correspondences resolved from geometry, not declared: which ring is which
brow, and which end is the outer one. The second matters more - backwards, the
tilt mirrors and worry renders as its own opposite, which reads as a directed
performance choice rather than a bug and would never be questioned. Which EDGE
is upper is deliberately left unresolved: it traverses the same ring the other
way, an even-odd fill has no winding, and both ends still land on fixed slots.
Also fixes a bug from the exposure work: the live render applied exposure to
the plate and the mouth but not to the eyes, so on 2s the preview and the
export disagreed. A preview that disagrees with the export is the one bug this
tool cannot afford. perfIndex now exists as a named thing so the two paths
cannot drift apart again.
91 -> 105 assertions. Ground truth on all four synthetic brow poses, tilt
separating worry from anger by sign, a blink not faking a raise, and a shared
dwell never emitting a half-raised brow.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:11:25 -04:00
|
|
|
EYE_R_LIDS, EYE_L_LIDS, BROW_A_RING, BROW_B_RING,
|
|
|
|
|
BROW_END_0, BROW_END_1 } from './landmarks.js';
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
import { fitSimilarity, applySim, procrustesMean, smoothTransforms, offsetRing } from './mathutil.js';
|
|
|
|
|
import { stabilize, toRasterRing, smoothContours, selectKeys, activeKey, shiftIndex,
|
Brows: traced ring, quantised raise
A brow at 320x200 is fourteen pixels wide and three tall. Its shape carries
almost nothing at that size; its height above the eye carries the expression,
and a brow raise is the most legible beat on a face. So the ring is traced and
the height is quantised - the split the eyes already got, where the lid is a
traced feature and the iris a quantised primitive.
The decomposition is the point. The traced ring already contains the real
height, so adding a quantised raise on top would move the brow twice. The
height is measured OUT of the ring, quantised, and put back, so the shape that
renders is his at a height that snaps between a few levels and holds.
Measured at both ends rather than as one number, because raise and tilt are
different expressions out of one mechanism: both ends up is surprise, inner up
alone is worry, inner down is anger. They share a dwell - the gaze quantiser,
renamed quantizeSnap now that it has two callers - so the brow hits its pose in
one frame instead of crawling into it with one end arriving before the other.
Measured against the eye's corner midpoint, never its lid. Same trap the gaze
origin has and worth avoiding twice: brows and lids move together constantly,
so a brow that jumped on every blink would read as a tic. Rest pose from the
take median rather than the neutral frame, for the reason gaze learned the hard
way - that frame is picked by minimum mouth aperture and says nothing about the
brows.
Two correspondences resolved from geometry, not declared: which ring is which
brow, and which end is the outer one. The second matters more - backwards, the
tilt mirrors and worry renders as its own opposite, which reads as a directed
performance choice rather than a bug and would never be questioned. Which EDGE
is upper is deliberately left unresolved: it traverses the same ring the other
way, an even-odd fill has no winding, and both ends still land on fixed slots.
Also fixes a bug from the exposure work: the live render applied exposure to
the plate and the mouth but not to the eyes, so on 2s the preview and the
export disagreed. A preview that disagrees with the export is the one bug this
tool cannot afford. perfIndex now exists as a named thing so the two paths
cannot drift apart again.
91 -> 105 assertions. Ground truth on all four synthetic brow poses, tilt
separating worry from anger by sign, a blink not faking a raise, and a shared
dwell never emitting a half-raised brow.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:11:25 -04:00
|
|
|
exposeIndex, eyeSignals, pairIrises, gazeOrigin, quantizeSnap, resolveBlink,
|
|
|
|
|
browSignals, pairBrows } from './pipeline.js';
|
2026-09-24 14:38:07 -04:00
|
|
|
import { IndexedRaster, hexToRgb } from './raster.js';
|
|
|
|
|
import { writeTake } from './take.js';
|
Teeth as an extracted blob contour, not a clipped band
The band filled the mouth because a band is the wrong reduction: the bright
region is a blob, and reading it as "everything above a line" throws the shape
away.
Extracting a contour reintroduces the vertex-correspondence problem that made
me avoid it, but for a blob there is a way out. Radial sampling from the
centroid along N fixed directions makes vertex k always mean "the extent in
direction k": correspondence holds by construction, the count is fixed, and
temporal smoothing cannot reorder anything. It also yields a star-shaped
reduction, which suits flat colour.
Tongue rejection, which the band had no way to express:
- pixels red relative to their own brightness are dropped (teeth are neutral)
- component choice is biased toward the top of the cavity, since area alone
picks the tongue when the mouth is wide
- separate inner and outer controls: cavity erode pulls the sampled region off
the lip edge, blob grow/erode resizes the found blob
Also: a knob wired in app.js but missing from index.html threw during wiring
and left a blank page with nothing useful in the console - which is exactly
what happened to teethDwell in the previous commit. el() now names the missing
id, and window.onerror surfaces it in the status line.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:24:52 -04:00
|
|
|
import { otsuForTest, scaleRing } from './interior.js';
|
2026-09-24 14:38:07 -04:00
|
|
|
import { synthDense } from './synth.js';
|
|
|
|
|
|
|
|
|
|
const results = [];
|
|
|
|
|
const ok = (name, cond, detail = '') => results.push({ name, pass: !!cond, detail });
|
|
|
|
|
|
|
|
|
|
/* ---- geometry helpers ---- */
|
|
|
|
|
|
|
|
|
|
function segmentsCross(a, b, c, d) {
|
|
|
|
|
const o = (p, q, r) => Math.sign((q.x - p.x) * (r.y - p.y) - (q.y - p.y) * (r.x - p.x));
|
|
|
|
|
const o1 = o(a, b, c), o2 = o(a, b, d), o3 = o(c, d, a), o4 = o(c, d, b);
|
|
|
|
|
return o1 !== o2 && o3 !== o4 && o1 !== 0 && o2 !== 0 && o3 !== 0 && o4 !== 0;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
// A closed ring is simple if no pair of non-adjacent edges crosses.
|
|
|
|
|
function ringSelfIntersections(pts) {
|
|
|
|
|
const n = pts.length, hits = [];
|
|
|
|
|
for (let i = 0; i < n; i++) {
|
|
|
|
|
for (let j = i + 1; j < n; j++) {
|
|
|
|
|
if (i === j || (j + 1) % n === i || (i + 1) % n === j) continue;
|
|
|
|
|
if (segmentsCross(pts[i], pts[(i + 1) % n], pts[j], pts[(j + 1) % n])) hits.push([i, j]);
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
return hits;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
const spreadX = (frames, slot) => {
|
|
|
|
|
const xs = frames.map((f) => f[slot].x);
|
|
|
|
|
return Math.max(...xs) - Math.min(...xs);
|
|
|
|
|
};
|
|
|
|
|
|
|
|
|
|
/* ---- the tests ---- */
|
|
|
|
|
|
2026-09-24 15:25:39 -04:00
|
|
|
// Cross-check every el('id') in app.js against the ids in index.html.
|
|
|
|
|
//
|
|
|
|
|
// This bug class has bitten twice: a knob wired in app.js but absent from the
|
|
|
|
|
// markup throws during wiring, which aborts the rest of the module and leaves a
|
|
|
|
|
// blank page. The symptom ("nothing happens") points nowhere near the cause, so
|
|
|
|
|
// it is worth an automated check rather than vigilance.
|
|
|
|
|
export async function runWiring() {
|
|
|
|
|
const out = [];
|
|
|
|
|
try {
|
|
|
|
|
const [app, html] = await Promise.all([
|
|
|
|
|
fetch('./js/app.js').then((r) => r.text()),
|
|
|
|
|
fetch('./index.html').then((r) => r.text()),
|
|
|
|
|
]);
|
|
|
|
|
const ids = new Set([...app.matchAll(/\bel\(\s*['"]([\w-]+)['"]\s*\)/g)].map((m) => m[1]));
|
|
|
|
|
const list = app.match(/for \(const id of \[([\s\S]*?)\]\)/);
|
|
|
|
|
if (list) {
|
|
|
|
|
for (const m of list[1].matchAll(/'([\w]+)'/g)) { ids.add(m[1]); ids.add(m[1] + 'v'); }
|
|
|
|
|
}
|
2026-09-24 15:36:04 -04:00
|
|
|
// Duplicate keys in an object literal are silent in JS - the last one wins.
|
|
|
|
|
// In opts() that meant the teeth vertex slider was quietly driving the lip
|
|
|
|
|
// vertex count while the lip slider did nothing at all.
|
|
|
|
|
const lit = app.match(/const opts = \(\) => \(\{([\s\S]*?)\n\}\);/);
|
|
|
|
|
if (lit) {
|
|
|
|
|
const keys = [...lit[1].matchAll(/^\s*([A-Za-z_$][\w$]*)\s*:/gm)].map((m) => m[1]);
|
|
|
|
|
const dupes = keys.filter((k, i) => keys.indexOf(k) !== i);
|
|
|
|
|
out.push({ name: `opts() has no duplicate keys (${keys.length} checked)`,
|
|
|
|
|
pass: dupes.length === 0, detail: [...new Set(dupes)].join(', ') });
|
|
|
|
|
} else {
|
|
|
|
|
out.push({ name: 'opts() literal found for duplicate-key check', pass: false, detail: '' });
|
|
|
|
|
}
|
|
|
|
|
|
2026-09-24 15:25:39 -04:00
|
|
|
const have = new Set([...html.matchAll(/id="([\w-]+)"/g)].map((m) => m[1]));
|
|
|
|
|
const missing = [...ids].filter((i) => !have.has(i));
|
|
|
|
|
out.push({ name: `every el() id exists in index.html (${ids.size} checked)`,
|
|
|
|
|
pass: missing.length === 0, detail: missing.join(', ') });
|
|
|
|
|
const unused = [...have].filter((i) => !ids.has(i));
|
|
|
|
|
out.push({ name: 'no orphaned ids in index.html', pass: unused.length === 0,
|
|
|
|
|
detail: unused.join(', ') });
|
|
|
|
|
} catch (e) {
|
|
|
|
|
out.push({ name: 'wiring check ran', pass: false, detail: e.message });
|
|
|
|
|
}
|
|
|
|
|
return out;
|
|
|
|
|
}
|
|
|
|
|
|
2026-09-24 14:38:07 -04:00
|
|
|
export function run() {
|
|
|
|
|
results.length = 0;
|
|
|
|
|
|
|
|
|
|
// tables
|
|
|
|
|
ok('LIPS_OUTER has 20 distinct ids', new Set(LIPS_OUTER).size === 20);
|
|
|
|
|
ok('LIPS_INNER has 20 distinct ids', new Set(LIPS_INNER).size === 20);
|
|
|
|
|
ok('FACE_OVAL has 36 distinct ids', new Set(FACE_OVAL).size === 36);
|
|
|
|
|
ok('RIGID excludes every lip vertex',
|
|
|
|
|
!RIGID.some((i) => LIPS_OUTER.includes(i) || LIPS_INNER.includes(i)),
|
|
|
|
|
'a moving feature in the rigid set bleeds performance into stabilisation');
|
|
|
|
|
|
|
|
|
|
// subsampling preserves order and count at every budget
|
|
|
|
|
for (let n = 4; n <= 16; n += 2) {
|
|
|
|
|
const s = subsampleSlots(20, n);
|
|
|
|
|
const mono = s.every((v, i) => i === 0 || v > s[i - 1]);
|
|
|
|
|
ok(`subsampleSlots(20,${n}) is strictly increasing, n=${n}`, mono && s.length === n, s.join(','));
|
|
|
|
|
}
|
|
|
|
|
ok('subsampleRing agrees with subsampleSlots',
|
|
|
|
|
subsampleRing(LIPS_OUTER, 8).join(',') === subsampleSlots(20, 8).map((s) => LIPS_OUTER[s]).join(','));
|
|
|
|
|
|
|
|
|
|
const dense = synthDense(72);
|
|
|
|
|
|
|
|
|
|
// rings must be simple at EVERY vertex budget, on every frame
|
|
|
|
|
for (const [label, table] of [['outer', LIPS_OUTER], ['inner', LIPS_INNER]]) {
|
|
|
|
|
let worst = null;
|
|
|
|
|
for (let n = 4; n <= 16 && !worst; n += 2) {
|
|
|
|
|
const slots = subsampleSlots(table.length, n);
|
|
|
|
|
for (let f = 0; f < dense.length; f++) {
|
|
|
|
|
const pts = slots.map((s) => dense[f][table[s]]);
|
|
|
|
|
const hits = ringSelfIntersections(pts);
|
|
|
|
|
if (hits.length) { worst = `verts=${n} frame=${f} edges ${JSON.stringify(hits[0])}`; break; }
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
ok(`${label} ring is simple at every vertex budget`, !worst, worst || '');
|
|
|
|
|
}
|
|
|
|
|
|
Registered photo underlay as the plate reference
The generated face oval was never going to be good enough to draw from:
MediaPipe's face oval is the FACE boundary, cut at the hairline and excluding
hair, ears, jaw and neck, so it is an egg by construction. Segmentation would
give a real head outline but costs a 16MB model and per-frame inference for a
shape that gets replaced by a drawing anyway.
So the plate layer becomes switchable, and the useful modes are photographic:
the source frame mapped into raster space through the same transform chain the
contours go through. Registration is the whole point - the head sits still and
a drawing traced from the underlay is already aligned to the mouth. An
unregistered underlay would be decoration.
- underlay.js: pixel->raster affine (a general affine, since MediaPipe
normalises x by width and y by height), registered draw, palette posterise
- plate modes: photo / photo dim / posterized / oval / oval+photo / none, B cycles
- worksheet cells are registered composites rather than raw crops
- Save frame 4x writes a 1280x800 PNG to draw on
- selftest: FACE_OVAL simplicity, which was never asserted; a wrong ordering
there reads as a lumpy plate rather than an obvious bowtie
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 14:57:59 -04:00
|
|
|
// FACE_OVAL traversal: never checked before, and a wrong ordering here shows up
|
|
|
|
|
// as a lumpy plate rather than an obvious bowtie, so it needs asserting.
|
|
|
|
|
{
|
|
|
|
|
let bad = null;
|
|
|
|
|
for (let f = 0; f < dense.length && !bad; f++) {
|
|
|
|
|
const h = ringSelfIntersections(FACE_OVAL.map((i) => dense[f][i]));
|
|
|
|
|
if (h.length) bad = `frame ${f} edges ${JSON.stringify(h[0])}`;
|
|
|
|
|
}
|
|
|
|
|
ok('FACE_OVAL is a simple ring on every frame', !bad, bad || '');
|
|
|
|
|
}
|
|
|
|
|
|
2026-09-24 14:38:07 -04:00
|
|
|
// similarity fit recovers a known transform
|
|
|
|
|
const src = [{ x: 0, y: 0 }, { x: 1, y: 0 }, { x: 0, y: 1 }, { x: 2, y: 3 }];
|
|
|
|
|
const truth = { s: 1.7, theta: 0.6, tx: 4, ty: -2 };
|
|
|
|
|
const dst = src.map((p) => applySim(truth, p));
|
|
|
|
|
const got = fitSimilarity(src, dst);
|
|
|
|
|
ok('fitSimilarity recovers a known transform',
|
|
|
|
|
Math.abs(got.s - truth.s) < 1e-9 && Math.abs(got.theta - truth.theta) < 1e-9 &&
|
|
|
|
|
Math.abs(got.tx - truth.tx) < 1e-9 && Math.abs(got.ty - truth.ty) < 1e-9,
|
|
|
|
|
`s=${got.s.toFixed(6)} th=${got.theta.toFixed(6)}`);
|
|
|
|
|
|
|
|
|
|
// stabilisation: head motion out, mouth motion kept
|
2026-09-24 14:51:15 -04:00
|
|
|
const stab = stabilize(dense, 0);
|
2026-09-24 14:38:07 -04:00
|
|
|
const rawSpread = spreadX(dense, 133);
|
|
|
|
|
const stabSpread = (() => {
|
|
|
|
|
const xs = stab.eyes.map((e) => e[0].x);
|
|
|
|
|
return Math.max(...xs) - Math.min(...xs);
|
|
|
|
|
})();
|
|
|
|
|
ok('stabilisation removes >90% of head translation',
|
|
|
|
|
stabSpread < rawSpread * 0.1, `raw ${rawSpread.toFixed(4)} -> ${stabSpread.toFixed(4)}`);
|
|
|
|
|
const apRange = Math.max(...stab.aperture) - Math.min(...stab.aperture);
|
|
|
|
|
ok('stabilisation preserves mouth motion', apRange > 0.05, `aperture range ${apRange.toFixed(4)}`);
|
|
|
|
|
|
2026-09-24 15:03:12 -04:00
|
|
|
// ASPECT: a shape that is circular in PIXEL space must stay circular in raster
|
|
|
|
|
// space. MediaPipe normalises x by width and y by height, so for a portrait
|
|
|
|
|
// frame equal normalised numbers are unequal pixel distances; feeding those
|
|
|
|
|
// straight through stretches everything horizontally by H/W. This asserts the
|
|
|
|
|
// isotropic conversion, and fails at ~1.78 for a 1080x1920 clip without it.
|
|
|
|
|
for (const [W, H] of [[1080, 1920], [1920, 1080], [640, 640]]) {
|
|
|
|
|
const aspect = W / H;
|
|
|
|
|
const N = 24, cx = 0.5, cy = 0.5, rPx = 200;
|
|
|
|
|
// a true circle of radius rPx, expressed in MediaPipe normalised coords
|
|
|
|
|
const circleFrames = [];
|
|
|
|
|
for (let t = 0; t < 4; t++) {
|
|
|
|
|
const pts = new Array(478).fill(null).map(() => ({ x: 0.5, y: 0.5, z: 0 }));
|
|
|
|
|
RIGID.forEach((id, k) => {
|
|
|
|
|
const a = (k / RIGID.length) * Math.PI * 2;
|
|
|
|
|
pts[id] = { x: cx + (120 * Math.cos(a)) / W, y: cy + (120 * Math.sin(a)) / H, z: 0 };
|
|
|
|
|
});
|
|
|
|
|
LIPS_OUTER.forEach((id, k) => {
|
|
|
|
|
const a = -(k / LIPS_OUTER.length) * Math.PI * 2;
|
|
|
|
|
pts[id] = { x: cx + (rPx * Math.cos(a)) / W, y: cy + (rPx * Math.sin(a)) / H, z: 0 };
|
|
|
|
|
});
|
|
|
|
|
FACE_OVAL.forEach((id, k) => {
|
|
|
|
|
const a = -(k / FACE_OVAL.length) * Math.PI * 2;
|
|
|
|
|
pts[id] = { x: cx + (420 * Math.cos(a)) / W, y: cy + (420 * Math.sin(a)) / H, z: 0 };
|
|
|
|
|
});
|
|
|
|
|
circleFrames.push(pts);
|
|
|
|
|
}
|
|
|
|
|
const st2 = stabilize(circleFrames, 0, aspect);
|
|
|
|
|
const ring = toRasterRing(st2.outer[0], LIPS_OUTER, 16, (p) => p);
|
|
|
|
|
const xs = ring.map((p) => p.x), ys = ring.map((p) => p.y);
|
|
|
|
|
const ratio = (Math.max(...xs) - Math.min(...xs)) / (Math.max(...ys) - Math.min(...ys));
|
|
|
|
|
ok(`circle stays circular at ${W}x${H}`, Math.abs(ratio - 1) < 0.02,
|
|
|
|
|
`w/h ratio ${ratio.toFixed(4)}`);
|
|
|
|
|
}
|
|
|
|
|
|
2026-09-24 14:38:07 -04:00
|
|
|
// key selection
|
|
|
|
|
const xf = (p) => ({ x: p.x * 320, y: p.y * 200 });
|
|
|
|
|
const shapes = stab.outer.map((r) => toRasterRing(r, LIPS_OUTER, 8, xf));
|
|
|
|
|
const sel = selectKeys(shapes, { minHold: 2, distThresh: 0.6, velSmooth: 3, exposure: 2 });
|
|
|
|
|
ok('keys are strictly increasing in f', sel.keys.every((k, i) => i === 0 || k.f > sel.keys[i - 1].f));
|
|
|
|
|
ok('keys respect the minimum hold',
|
|
|
|
|
sel.keys.every((k, i) => i === 0 || k.src - sel.keys[i - 1].src >= 2));
|
|
|
|
|
ok('keys land on the exposure grid', sel.keys.every((k) => k.f % 2 === 0));
|
|
|
|
|
ok('selection reduces candidates', sel.keys.length < sel.candidates.length,
|
|
|
|
|
`${sel.candidates.length} candidates -> ${sel.keys.length} keys`);
|
|
|
|
|
ok('first key is frame 0', sel.keys[0].f === 0);
|
|
|
|
|
ok('activeKey holds between keys',
|
|
|
|
|
activeKey(sel.keys, sel.keys[1].f - 1).f === sel.keys[0].f);
|
|
|
|
|
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
// exposure: rip dense, choose the timing here. On 2s every odd frame must
|
|
|
|
|
// reuse the even frame's pose, and the grid must never read from the future -
|
|
|
|
|
// that direction is the lead, which is a different control for a reason.
|
|
|
|
|
ok('exposure 1 is identity', [0, 1, 7, 71].every((f) => exposeIndex(f, 1) === f));
|
|
|
|
|
ok('on 2s holds each pose for two frames',
|
|
|
|
|
[0, 1, 2, 3, 4, 5].map((f) => exposeIndex(f, 2)).join(',') === '0,0,2,2,4,4');
|
|
|
|
|
ok('on 3s holds each pose for three frames',
|
|
|
|
|
[0, 1, 2, 3, 4, 5, 6].map((f) => exposeIndex(f, 3)).join(',') === '0,0,0,3,3,3,6');
|
|
|
|
|
ok('exposure never reads a pose from the future',
|
|
|
|
|
[0, 1, 2, 3, 4, 5, 6, 7].every((f) => exposeIndex(f, 3) <= f));
|
|
|
|
|
{
|
|
|
|
|
// Exposure then lead, in that order: the picture must change on the grid
|
|
|
|
|
// beats and carry a pose shifted by whole frames of the original track.
|
|
|
|
|
const N = 72, at = (f) => shiftIndex(exposeIndex(f, 2), 1, N);
|
|
|
|
|
ok('exposure and lead compose without moving the beats',
|
|
|
|
|
at(0) === 1 && at(1) === 1 && at(2) === 3 && at(3) === 3,
|
|
|
|
|
[0, 1, 2, 3].map(at).join(','));
|
|
|
|
|
}
|
|
|
|
|
|
2026-09-24 15:38:02 -04:00
|
|
|
// mouth lead: a shift that "feels like it does nothing" is indistinguishable
|
|
|
|
|
// from one that does nothing, so assert the arithmetic directly.
|
|
|
|
|
ok('lead 0 is identity', [0, 5, 71].every((f) => shiftIndex(f, 0, 72) === f));
|
|
|
|
|
ok('positive lead moves the source frame forward', shiftIndex(10, 2, 72) === 12);
|
|
|
|
|
ok('negative lead moves it back', shiftIndex(10, -3, 72) === 7);
|
|
|
|
|
ok('lead clamps at the start', shiftIndex(1, -6, 72) === 0);
|
|
|
|
|
ok('lead clamps at the end', shiftIndex(70, 6, 72) === 71);
|
|
|
|
|
{
|
|
|
|
|
// ...and that it selects different POSES, not merely different indices.
|
|
|
|
|
// Checked across the whole track rather than at one pair: synthetic poses
|
|
|
|
|
// hold for nine-frame beats, so any single pair can legitimately be
|
|
|
|
|
// identical while the shift works perfectly.
|
|
|
|
|
const N = shapes.length;
|
|
|
|
|
let moved = 0, total = 0;
|
|
|
|
|
for (let f = 0; f < N; f++) {
|
|
|
|
|
const a = shapes[shiftIndex(f, 0, N)], b = shapes[shiftIndex(f, 3, N)];
|
|
|
|
|
let d = 0;
|
|
|
|
|
for (let i = 0; i < a.length; i++) d += Math.hypot(a[i].x - b[i].x, a[i].y - b[i].y);
|
|
|
|
|
total++;
|
|
|
|
|
if (d / a.length > 0.5) moved++;
|
|
|
|
|
}
|
|
|
|
|
ok('a lead of 3 changes the pose on a good share of frames', moved / total > 0.2,
|
|
|
|
|
`${moved}/${total} frames differ`);
|
|
|
|
|
}
|
|
|
|
|
|
2026-09-24 14:38:07 -04:00
|
|
|
// rasteriser: indexed, hard-edged, no blending
|
|
|
|
|
const r = new IndexedRaster(64, 48);
|
|
|
|
|
r.clear(0);
|
|
|
|
|
r.fillPoly([{ x: 8, y: 8 }, { x: 56, y: 8 }, { x: 56, y: 40 }, { x: 8, y: 40 }], 2);
|
|
|
|
|
const present = new Set(r.buf);
|
|
|
|
|
ok('raster contains only written indices', present.size === 2 && present.has(0) && present.has(2),
|
|
|
|
|
`indices ${[...present].join(',')}`);
|
|
|
|
|
let count = 0;
|
|
|
|
|
for (const v of r.buf) if (v === 2) count++;
|
|
|
|
|
ok('axis-aligned rect fills the exact pixel count', count === 48 * 32, `${count} vs ${48 * 32}`);
|
|
|
|
|
const pal = ['#000000', '#ffffff', '#ff8800'];
|
|
|
|
|
const img = r.toImageData(pal, 2);
|
|
|
|
|
const seen = new Set();
|
|
|
|
|
for (let i = 0; i < img.data.length; i += 4) {
|
|
|
|
|
seen.add(`${img.data[i]},${img.data[i + 1]},${img.data[i + 2]}`);
|
|
|
|
|
}
|
|
|
|
|
const allowed = new Set(pal.map((h) => hexToRgb(h).join(',')));
|
|
|
|
|
ok('palette expansion introduces no intermediate colours',
|
|
|
|
|
[...seen].every((c) => allowed.has(c)), `${seen.size} distinct colours`);
|
|
|
|
|
|
Teeth as an extracted blob contour, not a clipped band
The band filled the mouth because a band is the wrong reduction: the bright
region is a blob, and reading it as "everything above a line" throws the shape
away.
Extracting a contour reintroduces the vertex-correspondence problem that made
me avoid it, but for a blob there is a way out. Radial sampling from the
centroid along N fixed directions makes vertex k always mean "the extent in
direction k": correspondence holds by construction, the count is fixed, and
temporal smoothing cannot reorder anything. It also yields a star-shaped
reduction, which suits flat colour.
Tongue rejection, which the band had no way to express:
- pixels red relative to their own brightness are dropped (teeth are neutral)
- component choice is biased toward the top of the cavity, since area alone
picks the tongue when the mouth is wide
- separate inner and outer controls: cavity erode pulls the sampled region off
the lip edge, blob grow/erode resizes the found blob
Also: a knob wired in app.js but missing from index.html threw during wiring
and left a blank page with nothing useful in the console - which is exactly
what happened to teethDwell in the previous commit. el() now names the missing
id, and window.onerror surfaces it in the status line.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:24:52 -04:00
|
|
|
// scaleRing is what pulls the sampled region in from MediaPipe's inner lip
|
|
|
|
|
// landmarks, which sit slightly outside the real opening.
|
2026-09-24 15:12:09 -04:00
|
|
|
{
|
|
|
|
|
const ring = [{ x: 0, y: 0 }, { x: 10, y: 0 }, { x: 10, y: 10 }, { x: 0, y: 10 }];
|
Teeth as an extracted blob contour, not a clipped band
The band filled the mouth because a band is the wrong reduction: the bright
region is a blob, and reading it as "everything above a line" throws the shape
away.
Extracting a contour reintroduces the vertex-correspondence problem that made
me avoid it, but for a blob there is a way out. Radial sampling from the
centroid along N fixed directions makes vertex k always mean "the extent in
direction k": correspondence holds by construction, the count is fixed, and
temporal smoothing cannot reorder anything. It also yields a star-shaped
reduction, which suits flat colour.
Tongue rejection, which the band had no way to express:
- pixels red relative to their own brightness are dropped (teeth are neutral)
- component choice is biased toward the top of the cavity, since area alone
picks the tongue when the mouth is wide
- separate inner and outer controls: cavity erode pulls the sampled region off
the lip edge, blob grow/erode resizes the found blob
Also: a knob wired in app.js but missing from index.html threw during wiring
and left a blank page with nothing useful in the console - which is exactly
what happened to teethDwell in the previous commit. el() now names the missing
id, and window.onerror surfaces it in the status line.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:24:52 -04:00
|
|
|
const small = scaleRing(ring, 0.5);
|
|
|
|
|
const w = Math.max(...small.map((p) => p.x)) - Math.min(...small.map((p) => p.x));
|
|
|
|
|
ok('scaleRing(0.5) halves the extent', Math.abs(w - 5) < 1e-9, `width ${w}`);
|
|
|
|
|
const same = scaleRing(ring, 1);
|
|
|
|
|
ok('scaleRing(1) is identity', same.every((p, i) => Math.abs(p.x - ring[i].x) < 1e-9));
|
|
|
|
|
let cx = 0; for (const p of small) cx += p.x;
|
|
|
|
|
ok('scaleRing keeps the centroid', Math.abs(cx / 4 - 5) < 1e-9);
|
2026-09-24 15:12:09 -04:00
|
|
|
}
|
|
|
|
|
|
Fix teeth band filling the whole mouth
Three causes, all of them mine:
Otsu always returns a split, including on a homogeneous region - given a dark
cavity with no teeth it invents a threshold and calls half the pixels bright.
The gate is now the separation between the two class means, which is the only
thing that says whether the split means anything. Coverage was the wrong
signal: it is high both when the mouth is full of teeth and when the region is
uniformly dark and Otsu has split noise.
The row scan tracked the last qualifying row anywhere rather than where the run
from the top stops, so one bright row near the bottom - a lit lower lip inside
the ring - pushed the line to full height. It now breaks at the first failing
row once the run has started.
MediaPipe's inner lip landmarks sit slightly outside the real opening, so the
sampled region included lip pixels, which are bright and sit exactly at the
boundary where they do most damage. The ring is now eroded toward its centroid
before sampling, with the amount exposed as a knob.
Adds a diagnostic panel showing the sampled crop, pixels above threshold, and
the resolved line, because tuning this from numbers alone does not tell you
whether the region being measured is even the right region.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:14:49 -04:00
|
|
|
// Otsu on a uniform region must report near-zero class separation. It will
|
|
|
|
|
// still return a threshold - that is what Otsu does - so the separation is the
|
|
|
|
|
// only thing that distinguishes "found teeth" from "split noise in a dark
|
|
|
|
|
// mouth", which is what made the band fill the whole cavity.
|
|
|
|
|
{
|
|
|
|
|
const flat = new Uint32Array(256); flat[40] = 500;
|
|
|
|
|
const f = otsuForTest(flat, 500);
|
|
|
|
|
ok('uniform region yields ~no class separation',
|
|
|
|
|
Math.abs(f.mBright - f.mDark) / 255 < 0.02, `sep ${((f.mBright - f.mDark) / 255).toFixed(4)}`);
|
|
|
|
|
|
|
|
|
|
const noisy = new Uint32Array(256);
|
|
|
|
|
for (let i = 30; i <= 60; i++) noisy[i] = 20; // dark cavity, some spread
|
|
|
|
|
const nz = otsuForTest(noisy, 31 * 20);
|
|
|
|
|
ok('dark-but-noisy region stays below a sane gate',
|
|
|
|
|
(nz.mBright - nz.mDark) / 255 < 0.14, `sep ${((nz.mBright - nz.mDark) / 255).toFixed(4)}`);
|
|
|
|
|
|
|
|
|
|
const teeth = new Uint32Array(256);
|
|
|
|
|
for (let i = 20; i <= 45; i++) teeth[i] = 40; // cavity
|
|
|
|
|
for (let i = 180; i <= 220; i++) teeth[i] = 30; // teeth
|
|
|
|
|
const tt = otsuForTest(teeth, 26 * 40 + 41 * 30);
|
|
|
|
|
ok('real bright/dark split clears the gate',
|
|
|
|
|
(tt.mBright - tt.mDark) / 255 > 0.4, `sep ${((tt.mBright - tt.mDark) / 255).toFixed(4)}`);
|
|
|
|
|
}
|
|
|
|
|
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
/* ---- eyes ---- */
|
|
|
|
|
|
|
|
|
|
ok('eye rings have 16 distinct ids each',
|
|
|
|
|
new Set(EYE_R_RING).size === 16 && new Set(EYE_L_RING).size === 16);
|
|
|
|
|
ok('the two eye rings share no landmark',
|
|
|
|
|
!EYE_R_RING.some((i) => EYE_L_RING.includes(i)));
|
|
|
|
|
|
|
|
|
|
// The cardinal contract, asserted rather than trusted: on a 16-slot ring the
|
|
|
|
|
// quarter slots must be the four anatomical cardinals, which is what makes
|
|
|
|
|
// every even vertex budget land on real landmarks instead of between them.
|
|
|
|
|
ok('eye ring slot 0/4/8/12 are outer, upper, inner, lower',
|
|
|
|
|
EYE_R_RING[0] === EYE_R_CORNERS[0] && EYE_R_RING[8] === EYE_R_CORNERS[1] &&
|
|
|
|
|
EYE_R_RING[4] === EYE_R_LIDS[0] && EYE_R_RING[12] === EYE_R_LIDS[1] &&
|
|
|
|
|
EYE_L_RING[0] === EYE_L_CORNERS[0] && EYE_L_RING[8] === EYE_L_CORNERS[1] &&
|
|
|
|
|
EYE_L_RING[4] === EYE_L_LIDS[0] && EYE_L_RING[12] === EYE_L_LIDS[1]);
|
|
|
|
|
|
|
|
|
|
// The gaze origin and denominator are built from the eye corners, so if a
|
|
|
|
|
// corner were not rigid a blink could move it and fake a glance.
|
|
|
|
|
ok('every eye corner is a rigid landmark',
|
|
|
|
|
[...EYE_R_CORNERS, ...EYE_L_CORNERS].every((i) => RIGID.includes(i)));
|
|
|
|
|
|
|
|
|
|
// Same simplicity requirement as the lips, and for the same reason: a cut
|
|
|
|
|
// part with a self-intersecting ring renders as blocks meeting at corners.
|
|
|
|
|
// Checked on blink frames too, where the ring is nearly degenerate.
|
|
|
|
|
for (const [label, table] of [['right', EYE_R_RING], ['left', EYE_L_RING]]) {
|
|
|
|
|
let worst = null;
|
|
|
|
|
for (let n = 4; n <= 12 && !worst; n += 2) {
|
|
|
|
|
const slots = subsampleSlots(table.length, n);
|
|
|
|
|
for (let f = 0; f < dense.length; f++) {
|
|
|
|
|
const hits = ringSelfIntersections(slots.map((sl) => dense[f][table[sl]]));
|
|
|
|
|
if (hits.length) { worst = `verts=${n} frame=${f} edges ${JSON.stringify(hits[0])}`; break; }
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
ok(`${label} eye ring is simple at every vertex budget`, !worst, worst || '');
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
// offsetRing must grow by a FIXED amount and survive a degenerate ring - the
|
|
|
|
|
// shut eyelid is exactly the degenerate case, and it is the frame where the
|
|
|
|
|
// lash line is the entire drawing.
|
|
|
|
|
{
|
|
|
|
|
const sq = [{ x: -1, y: 0 }, { x: 0, y: -1 }, { x: 1, y: 0 }, { x: 0, y: 1 }];
|
|
|
|
|
const g = offsetRing(sq, 2);
|
|
|
|
|
ok('offsetRing pushes every vertex out by exactly d',
|
|
|
|
|
g.every((p, i) => Math.abs(Math.hypot(p.x, p.y) - (Math.hypot(sq[i].x, sq[i].y) + 2)) < 1e-9));
|
|
|
|
|
ok('offsetRing(0) is identity', offsetRing(sq, 0) === sq);
|
|
|
|
|
// A shut lid: a flat sliver. The offset must still open it into a band.
|
|
|
|
|
const shutLid = [{ x: -10, y: 0 }, { x: 0, y: -0.02 }, { x: 10, y: 0 }, { x: 0, y: 0.02 }];
|
|
|
|
|
const band = offsetRing(shutLid, 1.5);
|
|
|
|
|
const h = Math.max(...band.map((p) => p.y)) - Math.min(...band.map((p) => p.y));
|
|
|
|
|
ok('offsetRing gives a shut lid a visible lash band', h > 2.9, `height ${h.toFixed(3)}`);
|
|
|
|
|
ok('offsetRing keeps the shut lid simple', ringSelfIntersections(band).length === 0);
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
{
|
|
|
|
|
const stE = stabilize(dense, 2);
|
|
|
|
|
const sig = eyeSignals(stE);
|
|
|
|
|
ok('synthetic track carries iris landmarks', sig.hasIris);
|
|
|
|
|
|
|
|
|
|
// THE load-bearing eye assertion. The pairing is resolved from geometry
|
|
|
|
|
// rather than declared, so the test feeds a track built the OTHER way round
|
|
|
|
|
// and demands the resolver follow the data. A resolver only ever checked
|
|
|
|
|
// against the convention it was written for is checking nothing.
|
|
|
|
|
const pairA = pairIrises(stE);
|
|
|
|
|
const pairB = pairIrises(stabilize(synthDense(72, { swapIris: true }), 2));
|
|
|
|
|
ok('iris pairing is resolved from the data, not assumed',
|
|
|
|
|
pairA.right === 'irisA' && pairB.right === 'irisB',
|
|
|
|
|
`normal ${pairA.right}, swapped ${pairB.right}`);
|
|
|
|
|
|
|
|
|
|
// The eye must TRACK the face, not sit in a fixed socket. An earlier
|
|
|
|
|
// version pinned each eye to its corners' mean over the shot, which does
|
|
|
|
|
// kill the wobble but leaves the drawn eyes hanging still over a registered
|
|
|
|
|
// photo whose eyes are moving. Head-local is the same space the mouth and
|
|
|
|
|
// the underlay live in, so the eye moves with the head exactly as they do.
|
|
|
|
|
{
|
|
|
|
|
const spread = (arr, sel) => {
|
|
|
|
|
const v = arr.map(sel);
|
|
|
|
|
return Math.max(...v) - Math.min(...v);
|
|
|
|
|
};
|
|
|
|
|
const w = Math.hypot(stE.cornersR[0][0].x - stE.cornersR[0][1].x,
|
|
|
|
|
stE.cornersR[0][0].y - stE.cornersR[0][1].y);
|
|
|
|
|
const moves = Math.max(spread(stE.lidR, (r) => r[0].x), spread(stE.lidR, (r) => r[0].y));
|
|
|
|
|
ok('the eye stays in head-local space and tracks the face', moves / w > 0.02,
|
|
|
|
|
`corner travels ${(moves / w * 100).toFixed(1)}% of an eye width`);
|
|
|
|
|
|
|
|
|
|
// Subsampling a 16-slot ring to any even budget must keep the two corners
|
|
|
|
|
// at output indices 0 and n/2. That is what lets the socket be read back
|
|
|
|
|
// off the drawn polygon instead of measured separately, which is what
|
|
|
|
|
// stops the iris drifting relative to the eye it sits in.
|
|
|
|
|
let bad = null;
|
|
|
|
|
for (let n = 4; n <= 12; n += 2) {
|
|
|
|
|
const sl = subsampleSlots(16, n);
|
|
|
|
|
if (sl[0] !== 0 || sl[n / 2] !== 8) bad = `n=${n} -> ${sl.join(',')}`;
|
|
|
|
|
}
|
|
|
|
|
ok('the drawn lid ring carries its own corners at 0 and n/2', !bad, bad || '');
|
|
|
|
|
|
|
|
|
|
// The contour average is what removes the jitter, and it is the mouth's
|
|
|
|
|
// knob doing the mouth's job - no second mechanism for the eyes.
|
|
|
|
|
const ring = (rad) => smoothContours(
|
|
|
|
|
stE.lidR.map((r) => toRasterRing(r, EYE_R_RING, 8, (p) => ({ x: p.x * 600, y: p.y * 600 }))), rad);
|
|
|
|
|
const jitter = (rings) => {
|
|
|
|
|
let acc = 0;
|
|
|
|
|
for (let f = 1; f < rings.length; f++) {
|
|
|
|
|
const a = rings[f], b = rings[f - 1];
|
|
|
|
|
acc += Math.hypot((a[0].x + a[4].x) / 2 - (b[0].x + b[4].x) / 2,
|
|
|
|
|
(a[0].y + a[4].y) / 2 - (b[0].y + b[4].y) / 2);
|
|
|
|
|
}
|
|
|
|
|
return acc / (rings.length - 1);
|
|
|
|
|
};
|
|
|
|
|
ok('contour averaging steadies the eye without pinning it',
|
|
|
|
|
jitter(ring(1)) < jitter(ring(0)) * 0.8,
|
|
|
|
|
`${jitter(ring(0)).toFixed(3)} -> ${jitter(ring(1)).toFixed(3)} px/frame`);
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
// Blink: synth shuts the lids for exactly one frame every 19.
|
|
|
|
|
const lo = Math.min(...sig.openR), hi = Math.max(...sig.openR);
|
|
|
|
|
ok('openness collapses on a blink and not otherwise', lo < hi * 0.2,
|
|
|
|
|
`${lo.toFixed(3)} .. ${hi.toFixed(3)}`);
|
|
|
|
|
|
|
|
|
|
const shut = resolveBlink(sig.openR, { cut: hi * 0.3, dwell: 0, hold: 3 });
|
|
|
|
|
const runs = [];
|
|
|
|
|
for (let f = 0; f < shut.length; f++) if (shut[f] && !shut[f - 1]) runs.push(f);
|
|
|
|
|
const lens = runs.map((a) => { let n = 0; while (shut[a + n]) n++; return n; });
|
|
|
|
|
ok('blinks are found', runs.length >= 3, `${runs.length} runs at ${runs.join(',')}`);
|
|
|
|
|
// The knob that is not like the teeth: a one-frame blink reads as a dropped
|
|
|
|
|
// frame, so `hold` must stretch it into something legible.
|
|
|
|
|
ok('a one-frame blink is held to the minimum length',
|
|
|
|
|
lens.every((n) => n >= 3), `run lengths ${lens.join(',')}`);
|
|
|
|
|
ok('a shorter hold leaves the blink shorter',
|
|
|
|
|
resolveBlink(sig.openR, { cut: hi * 0.3, dwell: 0, hold: 1 }).filter(Boolean).length <
|
|
|
|
|
shut.filter(Boolean).length);
|
|
|
|
|
|
|
|
|
|
// Gaze, against ground truth: synth commands +0.16 eye widths at f12 and
|
|
|
|
|
// -0.16 at f23, holding each for eleven frames.
|
|
|
|
|
const org = gazeOrigin(sig.gazeRaw, 'neutral', 0);
|
|
|
|
|
const gx = (f) => (sig.gazeRaw[f].x - org.x);
|
|
|
|
|
ok('gaze recovers the commanded direction',
|
|
|
|
|
gx(12) > 0.12 && gx(12) < 0.20 && gx(23) < -0.12 && gx(23) > -0.20,
|
|
|
|
|
`f12 ${gx(12).toFixed(3)}, f23 ${gx(23).toFixed(3)}`);
|
|
|
|
|
|
|
|
|
|
// Measuring gaze against the lid centroid instead of the corner midpoint
|
|
|
|
|
// would drag the iris down on every blink and fake a glance at the floor,
|
|
|
|
|
// on exactly the frames where the eye is most conspicuous.
|
|
|
|
|
const gy = (f) => (sig.gazeRaw[f].y - org.y);
|
|
|
|
|
ok('a blink does not fake a change of gaze',
|
|
|
|
|
Math.abs(gy(19) - gy(18)) < 0.02, `f18 ${gy(18).toFixed(4)} -> f19 ${gy(19).toFixed(4)}`);
|
|
|
|
|
|
|
|
|
|
// Quantisation is what turns drift into saccades: four commanded
|
|
|
|
|
// fixations must come back as a handful of cells, not one per frame.
|
|
|
|
|
// The origin re-points the whole performance, so a wrong one does not bias
|
|
|
|
|
// the gaze slightly - it makes the character look the other way. The median
|
|
|
|
|
// must sit inside the range it summarises; the neutral-frame origin need
|
|
|
|
|
// not, which is exactly the failure mode it has on footage with no
|
|
|
|
|
// deliberate neutral at the top.
|
|
|
|
|
{
|
|
|
|
|
const med = gazeOrigin(sig.gazeRaw, 'median');
|
|
|
|
|
const xs = sig.gazeRaw.map((g) => g.x);
|
|
|
|
|
ok('the median origin lies inside the take\'s own gaze range',
|
|
|
|
|
med.x > Math.min(...xs) && med.x < Math.max(...xs),
|
|
|
|
|
`${med.x.toFixed(3)} in ${Math.min(...xs).toFixed(3)}..${Math.max(...xs).toFixed(3)}`);
|
|
|
|
|
// Synth looks left as much as right, so the rest point is near zero.
|
|
|
|
|
ok('the median origin finds the rest point, not a glance',
|
|
|
|
|
Math.abs(med.x) < 0.08, `median x ${med.x.toFixed(3)}`);
|
|
|
|
|
ok('the two origins actually differ, so the toggle is a real A/B',
|
|
|
|
|
Math.abs(med.x - gazeOrigin(sig.gazeRaw, 'neutral', 12).x) > 0.02);
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
const px = sig.gazeRaw.map((g) => ({ x: (g.x - org.x) * 30, y: (g.y - org.y) * 30 }));
|
|
|
|
|
const cells = (a) => new Set(a.map((g) => `${g.x},${g.y}`)).size;
|
|
|
|
|
ok('quantisation collapses drift into a few fixations',
|
Brows: traced ring, quantised raise
A brow at 320x200 is fourteen pixels wide and three tall. Its shape carries
almost nothing at that size; its height above the eye carries the expression,
and a brow raise is the most legible beat on a face. So the ring is traced and
the height is quantised - the split the eyes already got, where the lid is a
traced feature and the iris a quantised primitive.
The decomposition is the point. The traced ring already contains the real
height, so adding a quantised raise on top would move the brow twice. The
height is measured OUT of the ring, quantised, and put back, so the shape that
renders is his at a height that snaps between a few levels and holds.
Measured at both ends rather than as one number, because raise and tilt are
different expressions out of one mechanism: both ends up is surprise, inner up
alone is worry, inner down is anger. They share a dwell - the gaze quantiser,
renamed quantizeSnap now that it has two callers - so the brow hits its pose in
one frame instead of crawling into it with one end arriving before the other.
Measured against the eye's corner midpoint, never its lid. Same trap the gaze
origin has and worth avoiding twice: brows and lids move together constantly,
so a brow that jumped on every blink would read as a tic. Rest pose from the
take median rather than the neutral frame, for the reason gaze learned the hard
way - that frame is picked by minimum mouth aperture and says nothing about the
brows.
Two correspondences resolved from geometry, not declared: which ring is which
brow, and which end is the outer one. The second matters more - backwards, the
tilt mirrors and worry renders as its own opposite, which reads as a directed
performance choice rather than a bug and would never be questioned. Which EDGE
is upper is deliberately left unresolved: it traverses the same ring the other
way, an even-odd fill has no winding, and both ends still land on fixed slots.
Also fixes a bug from the exposure work: the live render applied exposure to
the plate and the mouth but not to the eyes, so on 2s the preview and the
export disagreed. A preview that disagrees with the export is the one bug this
tool cannot afford. perfIndex now exists as a named thing so the two paths
cannot drift apart again.
91 -> 105 assertions. Ground truth on all four synthetic brow poses, tilt
separating worry from anger by sign, a blink not faking a raise, and a shared
dwell never emitting a half-raised brow.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:11:25 -04:00
|
|
|
cells(quantizeSnap(px, 2, 2)) <= 6 && cells(px) > 40,
|
|
|
|
|
`${cells(px)} raw -> ${cells(quantizeSnap(px, 2, 2))} cells`);
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
ok('gaze step 0 leaves the track untouched',
|
Brows: traced ring, quantised raise
A brow at 320x200 is fourteen pixels wide and three tall. Its shape carries
almost nothing at that size; its height above the eye carries the expression,
and a brow raise is the most legible beat on a face. So the ring is traced and
the height is quantised - the split the eyes already got, where the lid is a
traced feature and the iris a quantised primitive.
The decomposition is the point. The traced ring already contains the real
height, so adding a quantised raise on top would move the brow twice. The
height is measured OUT of the ring, quantised, and put back, so the shape that
renders is his at a height that snaps between a few levels and holds.
Measured at both ends rather than as one number, because raise and tilt are
different expressions out of one mechanism: both ends up is surprise, inner up
alone is worry, inner down is anger. They share a dwell - the gaze quantiser,
renamed quantizeSnap now that it has two callers - so the brow hits its pose in
one frame instead of crawling into it with one end arriving before the other.
Measured against the eye's corner midpoint, never its lid. Same trap the gaze
origin has and worth avoiding twice: brows and lids move together constantly,
so a brow that jumped on every blink would read as a tic. Rest pose from the
take median rather than the neutral frame, for the reason gaze learned the hard
way - that frame is picked by minimum mouth aperture and says nothing about the
brows.
Two correspondences resolved from geometry, not declared: which ring is which
brow, and which end is the outer one. The second matters more - backwards, the
tilt mirrors and worry renders as its own opposite, which reads as a directed
performance choice rather than a bug and would never be questioned. Which EDGE
is upper is deliberately left unresolved: it traverses the same ring the other
way, an even-odd fill has no winding, and both ends still land on fixed slots.
Also fixes a bug from the exposure work: the live render applied exposure to
the plate and the mouth but not to the eyes, so on 2s the preview and the
export disagreed. A preview that disagrees with the export is the one bug this
tool cannot afford. perfIndex now exists as a named thing so the two paths
cannot drift apart again.
91 -> 105 assertions. Ground truth on all four synthetic brow poses, tilt
separating worry from anger by sign, a blink not faking a raise, and a shared
dwell never emitting a half-raised brow.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:11:25 -04:00
|
|
|
quantizeSnap(px, 0, 2).every((g, i) => g.x === px[i].x && g.y === px[i].y));
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
ok('quantised values land on the grid',
|
Brows: traced ring, quantised raise
A brow at 320x200 is fourteen pixels wide and three tall. Its shape carries
almost nothing at that size; its height above the eye carries the expression,
and a brow raise is the most legible beat on a face. So the ring is traced and
the height is quantised - the split the eyes already got, where the lid is a
traced feature and the iris a quantised primitive.
The decomposition is the point. The traced ring already contains the real
height, so adding a quantised raise on top would move the brow twice. The
height is measured OUT of the ring, quantised, and put back, so the shape that
renders is his at a height that snaps between a few levels and holds.
Measured at both ends rather than as one number, because raise and tilt are
different expressions out of one mechanism: both ends up is surprise, inner up
alone is worry, inner down is anger. They share a dwell - the gaze quantiser,
renamed quantizeSnap now that it has two callers - so the brow hits its pose in
one frame instead of crawling into it with one end arriving before the other.
Measured against the eye's corner midpoint, never its lid. Same trap the gaze
origin has and worth avoiding twice: brows and lids move together constantly,
so a brow that jumped on every blink would read as a tic. Rest pose from the
take median rather than the neutral frame, for the reason gaze learned the hard
way - that frame is picked by minimum mouth aperture and says nothing about the
brows.
Two correspondences resolved from geometry, not declared: which ring is which
brow, and which end is the outer one. The second matters more - backwards, the
tilt mirrors and worry renders as its own opposite, which reads as a directed
performance choice rather than a bug and would never be questioned. Which EDGE
is upper is deliberately left unresolved: it traverses the same ring the other
way, an even-odd fill has no winding, and both ends still land on fixed slots.
Also fixes a bug from the exposure work: the live render applied exposure to
the plate and the mouth but not to the eyes, so on 2s the preview and the
export disagreed. A preview that disagrees with the export is the one bug this
tool cannot afford. perfIndex now exists as a named thing so the two paths
cannot drift apart again.
91 -> 105 assertions. Ground truth on all four synthetic brow poses, tilt
separating worry from anger by sign, a blink not faking a raise, and a shared
dwell never emitting a half-raised brow.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:11:25 -04:00
|
|
|
quantizeSnap(px, 2, 0).every((g) => Math.abs(g.x % 2) < 1e-9 && Math.abs(g.y % 2) < 1e-9));
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
// A one-frame excursion is noise; the dwell must swallow it.
|
|
|
|
|
{
|
|
|
|
|
const spike = [{ x: 0, y: 0 }, { x: 0, y: 0 }, { x: 4, y: 0 }, { x: 0, y: 0 }, { x: 0, y: 0 }];
|
|
|
|
|
ok('the dwell suppresses a one-frame gaze spike',
|
Brows: traced ring, quantised raise
A brow at 320x200 is fourteen pixels wide and three tall. Its shape carries
almost nothing at that size; its height above the eye carries the expression,
and a brow raise is the most legible beat on a face. So the ring is traced and
the height is quantised - the split the eyes already got, where the lid is a
traced feature and the iris a quantised primitive.
The decomposition is the point. The traced ring already contains the real
height, so adding a quantised raise on top would move the brow twice. The
height is measured OUT of the ring, quantised, and put back, so the shape that
renders is his at a height that snaps between a few levels and holds.
Measured at both ends rather than as one number, because raise and tilt are
different expressions out of one mechanism: both ends up is surprise, inner up
alone is worry, inner down is anger. They share a dwell - the gaze quantiser,
renamed quantizeSnap now that it has two callers - so the brow hits its pose in
one frame instead of crawling into it with one end arriving before the other.
Measured against the eye's corner midpoint, never its lid. Same trap the gaze
origin has and worth avoiding twice: brows and lids move together constantly,
so a brow that jumped on every blink would read as a tic. Rest pose from the
take median rather than the neutral frame, for the reason gaze learned the hard
way - that frame is picked by minimum mouth aperture and says nothing about the
brows.
Two correspondences resolved from geometry, not declared: which ring is which
brow, and which end is the outer one. The second matters more - backwards, the
tilt mirrors and worry renders as its own opposite, which reads as a directed
performance choice rather than a bug and would never be questioned. Which EDGE
is upper is deliberately left unresolved: it traverses the same ring the other
way, an even-odd fill has no winding, and both ends still land on fixed slots.
Also fixes a bug from the exposure work: the live render applied exposure to
the plate and the mouth but not to the eyes, so on 2s the preview and the
export disagreed. A preview that disagrees with the export is the one bug this
tool cannot afford. perfIndex now exists as a named thing so the two paths
cannot drift apart again.
91 -> 105 assertions. Ground truth on all four synthetic brow poses, tilt
separating worry from anger by sign, a blink not faking a raise, and a shared
dwell never emitting a half-raised brow.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:11:25 -04:00
|
|
|
quantizeSnap(spike, 2, 1).every((g) => g.x === 0));
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
ok('a sustained move still gets through',
|
Brows: traced ring, quantised raise
A brow at 320x200 is fourteen pixels wide and three tall. Its shape carries
almost nothing at that size; its height above the eye carries the expression,
and a brow raise is the most legible beat on a face. So the ring is traced and
the height is quantised - the split the eyes already got, where the lid is a
traced feature and the iris a quantised primitive.
The decomposition is the point. The traced ring already contains the real
height, so adding a quantised raise on top would move the brow twice. The
height is measured OUT of the ring, quantised, and put back, so the shape that
renders is his at a height that snaps between a few levels and holds.
Measured at both ends rather than as one number, because raise and tilt are
different expressions out of one mechanism: both ends up is surprise, inner up
alone is worry, inner down is anger. They share a dwell - the gaze quantiser,
renamed quantizeSnap now that it has two callers - so the brow hits its pose in
one frame instead of crawling into it with one end arriving before the other.
Measured against the eye's corner midpoint, never its lid. Same trap the gaze
origin has and worth avoiding twice: brows and lids move together constantly,
so a brow that jumped on every blink would read as a tic. Rest pose from the
take median rather than the neutral frame, for the reason gaze learned the hard
way - that frame is picked by minimum mouth aperture and says nothing about the
brows.
Two correspondences resolved from geometry, not declared: which ring is which
brow, and which end is the outer one. The second matters more - backwards, the
tilt mirrors and worry renders as its own opposite, which reads as a directed
performance choice rather than a bug and would never be questioned. Which EDGE
is upper is deliberately left unresolved: it traverses the same ring the other
way, an even-odd fill has no winding, and both ends still land on fixed slots.
Also fixes a bug from the exposure work: the live render applied exposure to
the plate and the mouth but not to the eyes, so on 2s the preview and the
export disagreed. A preview that disagrees with the export is the one bug this
tool cannot afford. perfIndex now exists as a named thing so the two paths
cannot drift apart again.
91 -> 105 assertions. Ground truth on all four synthetic brow poses, tilt
separating worry from anger by sign, a blink not faking a raise, and a shared
dwell never emitting a half-raised brow.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:11:25 -04:00
|
|
|
quantizeSnap([...spike, { x: 4, y: 0 }, { x: 4, y: 0 }, { x: 4, y: 0 }], 2, 1).pop().x === 4);
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
/* ---- brows ---- */
|
|
|
|
|
|
|
|
|
|
ok('brow rings have 10 distinct ids each',
|
|
|
|
|
new Set(BROW_A_RING).size === 10 && new Set(BROW_B_RING).size === 10);
|
|
|
|
|
ok('the brow rings share no landmark with each other, RIGID, or the lids',
|
|
|
|
|
!BROW_A_RING.some((i) => BROW_B_RING.includes(i)) &&
|
|
|
|
|
![...BROW_A_RING, ...BROW_B_RING].some((i) =>
|
|
|
|
|
RIGID.includes(i) || EYE_R_RING.includes(i) || EYE_L_RING.includes(i)),
|
|
|
|
|
'a brow in RIGID would bleed expression into the stabilisation');
|
|
|
|
|
|
|
|
|
|
{
|
|
|
|
|
let bad = null;
|
|
|
|
|
for (const [label, table] of [['A', BROW_A_RING], ['B', BROW_B_RING]]) {
|
|
|
|
|
for (let n = 4; n <= 10 && !bad; n += 2) {
|
|
|
|
|
const slots = subsampleSlots(table.length, n);
|
|
|
|
|
for (let f = 0; f < dense.length; f++) {
|
|
|
|
|
if (ringSelfIntersections(slots.map((sl) => dense[f][table[sl]])).length) {
|
|
|
|
|
bad = `${label} verts=${n} frame=${f}`; break;
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
ok('brow rings are simple at every vertex budget', !bad, bad || '');
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
{
|
|
|
|
|
const stB = stabilize(dense, 2);
|
|
|
|
|
const pr = pairBrows(stB);
|
|
|
|
|
ok('brow-to-eye pairing is resolved from geometry',
|
|
|
|
|
pr.right === 'browA' && pr.left === 'browB', JSON.stringify(pr));
|
|
|
|
|
// Getting this backwards mirrors the tilt, so inner-up "worried" renders as
|
|
|
|
|
// outer-up. That is a different expression, not a broken one, which is
|
|
|
|
|
// exactly why it needs an assertion rather than an eyeball.
|
|
|
|
|
ok('the outer end of the brow ring is resolved from geometry', pr.outerAtSlot0 === true);
|
|
|
|
|
|
|
|
|
|
// Both ends land on fixed slots whichever edge of the brow is on top, which
|
|
|
|
|
// is what lets the upper/lower ambiguity go unresolved without consequence.
|
|
|
|
|
ok('brow end slots are disjoint and cover both ends',
|
|
|
|
|
!BROW_END_0.some((i) => BROW_END_1.includes(i)) &&
|
|
|
|
|
BROW_END_0.length === 2 && BROW_END_1.length === 2);
|
|
|
|
|
|
|
|
|
|
// Ground truth: synth commands rest, surprise, worry and anger as heights
|
|
|
|
|
// above the eye centre in eye widths, holding each for thirteen frames.
|
|
|
|
|
const b = browSignals(stB);
|
|
|
|
|
const at = (f) => [b.R[f].x, b.R[f].y];
|
|
|
|
|
const near = (v, want) => Math.abs(v - want) < 0.02;
|
|
|
|
|
ok('brow raise recovers the commanded rest pose', near(at(0)[0], 0.30) && near(at(0)[1], 0.30),
|
|
|
|
|
at(0).map((v) => v.toFixed(3)).join(', '));
|
|
|
|
|
ok('brow raise recovers surprise - both ends up',
|
|
|
|
|
near(at(14)[0], 0.46) && near(at(14)[1], 0.46), at(14).map((v) => v.toFixed(3)).join(', '));
|
|
|
|
|
ok('brow raise recovers worry - inner end only',
|
|
|
|
|
near(at(27)[0], 0.30) && near(at(27)[1], 0.44), at(27).map((v) => v.toFixed(3)).join(', '));
|
|
|
|
|
ok('brow raise recovers anger - inner end down',
|
|
|
|
|
near(at(40)[0], 0.30) && near(at(40)[1], 0.18), at(40).map((v) => v.toFixed(3)).join(', '));
|
|
|
|
|
|
|
|
|
|
// Tilt must be a signed quantity that separates worry from anger. If the
|
|
|
|
|
// outer/inner resolution were mirrored these two would swap.
|
|
|
|
|
ok('tilt separates worry from anger by sign',
|
|
|
|
|
(at(27)[1] - at(27)[0]) > 0.08 && (at(40)[1] - at(40)[0]) < -0.08,
|
|
|
|
|
`worry ${(at(27)[1] - at(27)[0]).toFixed(3)}, anger ${(at(40)[1] - at(40)[0]).toFixed(3)}`);
|
|
|
|
|
|
|
|
|
|
// A blink must not read as a brow raise: the raise is measured against the
|
|
|
|
|
// eye's rigid corners, not its lid, which is the same trap the gaze origin
|
|
|
|
|
// has and worth avoiding twice.
|
|
|
|
|
const dR = Math.abs(b.R[19].x - b.R[18].x);
|
|
|
|
|
ok('a blink does not fake a brow raise', dR < 0.01, `f18 -> f19 delta ${dR.toFixed(4)}`);
|
|
|
|
|
|
|
|
|
|
// Quantisation: four sustained poses must come back as a handful of levels.
|
|
|
|
|
const rest = gazeOrigin(b.R, 'median');
|
|
|
|
|
const px = b.R.map((g) => ({ x: (g.x - rest.x) * 60, y: (g.y - rest.y) * 60 }));
|
|
|
|
|
const cells = (a) => new Set(a.map((g) => `${g.x},${g.y}`)).size;
|
|
|
|
|
ok('brow quantisation collapses drift into a few poses',
|
|
|
|
|
cells(quantizeSnap(px, 2, 2)) <= 6 && cells(px) > 20,
|
|
|
|
|
`${cells(px)} raw -> ${cells(quantizeSnap(px, 2, 2))} poses`);
|
|
|
|
|
// The dwell is SHARED across both channels, and that is the whole reason
|
|
|
|
|
// brows reuse the gaze quantiser rather than running two independent ones.
|
|
|
|
|
// Here the outer end moves one frame before the inner: with a shared dwell
|
|
|
|
|
// the half-raised pose (2,0) is transient and never commits, so the brow
|
|
|
|
|
// snaps once. Two independent dwells would emit it and the brow would crawl
|
|
|
|
|
// into position over two frames instead of hitting it.
|
|
|
|
|
{
|
|
|
|
|
const staggered = [
|
|
|
|
|
{ x: 0, y: 0 }, { x: 0, y: 0 }, { x: 2, y: 0 },
|
|
|
|
|
{ x: 2, y: 2 }, { x: 2, y: 2 }, { x: 2, y: 2 },
|
|
|
|
|
];
|
|
|
|
|
const out = quantizeSnap(staggered, 2, 1);
|
|
|
|
|
ok('a shared dwell never emits a half-raised brow',
|
|
|
|
|
!out.some((g) => g.x === 2 && g.y === 0),
|
|
|
|
|
out.map((g) => `${g.x},${g.y}`).join(' '));
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
// The iris is stencilled by the sclera and the pupil by the iris, which is
|
|
|
|
|
// what keeps both inside the lid at any gaze without clamping the gaze itself.
|
|
|
|
|
{
|
|
|
|
|
const rr = new IndexedRaster(40, 40);
|
|
|
|
|
rr.clear(0);
|
|
|
|
|
rr.fillPoly([{ x: 10, y: 10 }, { x: 30, y: 10 }, { x: 30, y: 20 }, { x: 10, y: 20 }], 1);
|
|
|
|
|
rr.fillDisc(28, 15, 9, 2, 1); // a disc reaching well past the "lid"
|
|
|
|
|
let spill = 0, inside = 0;
|
|
|
|
|
for (let y = 0; y < 40; y++) for (let x = 0; x < 40; x++) {
|
|
|
|
|
const v = rr.buf[y * 40 + x];
|
|
|
|
|
if (v !== 2) continue;
|
|
|
|
|
if (x >= 10 && x < 30 && y >= 10 && y < 20) inside++; else spill++;
|
|
|
|
|
}
|
|
|
|
|
ok('a stencilled disc cannot spill past its clip', spill === 0 && inside > 20,
|
|
|
|
|
`${inside} in, ${spill} out`);
|
|
|
|
|
rr.fillDisc(5, 35, 3, 3); // no stencil: writes freely
|
|
|
|
|
ok('an unstencilled disc still writes anywhere', rr.buf.includes(3));
|
|
|
|
|
|
|
|
|
|
// The stencil chain: pupil over iris over sclera. A pupil placed where the
|
|
|
|
|
// iris has already been cropped must be cropped the same way.
|
|
|
|
|
rr.fillRect(28, 15, 5, 4, 2);
|
|
|
|
|
let pSpill = 0;
|
|
|
|
|
for (let y = 0; y < 40; y++) for (let x = 0; x < 40; x++) {
|
|
|
|
|
if (rr.buf[y * 40 + x] === 4 && !(x >= 10 && x < 30 && y >= 10 && y < 20)) pSpill++;
|
|
|
|
|
}
|
|
|
|
|
ok('the pupil inherits the iris clip transitively', pSpill === 0);
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
// A square pupil is only worth having if it is the SAME square every frame:
|
|
|
|
|
// exactly its nominal size at any centre, or it breathes as the gaze moves.
|
|
|
|
|
{
|
|
|
|
|
const sizes = [];
|
|
|
|
|
for (const [cx, cy] of [[20, 20], [20.5, 20.5], [20.49, 19.51], [21, 20]]) {
|
|
|
|
|
const rr = new IndexedRaster(40, 40);
|
|
|
|
|
rr.clear(0);
|
|
|
|
|
rr.fillRect(cx, cy, 3, 1);
|
|
|
|
|
let n = 0, minX = 99, maxX = -1, minY = 99, maxY = -1;
|
|
|
|
|
for (let y = 0; y < 40; y++) for (let x = 0; x < 40; x++) {
|
|
|
|
|
if (rr.buf[y * 40 + x] !== 1) continue;
|
|
|
|
|
n++; minX = Math.min(minX, x); maxX = Math.max(maxX, x);
|
|
|
|
|
minY = Math.min(minY, y); maxY = Math.max(maxY, y);
|
|
|
|
|
}
|
|
|
|
|
sizes.push(`${maxX - minX + 1}x${maxY - minY + 1}:${n}`);
|
|
|
|
|
}
|
|
|
|
|
ok('a 3px pupil is 3x3 at every centre', sizes.every((v) => v === '3x3:9'), sizes.join(' '));
|
|
|
|
|
const rr = new IndexedRaster(40, 40);
|
|
|
|
|
rr.clear(0); rr.fillRect(20, 20, 0, 1);
|
|
|
|
|
ok('pupil size 0 draws nothing', !rr.buf.includes(1));
|
|
|
|
|
}
|
|
|
|
|
|
2026-09-24 14:38:07 -04:00
|
|
|
// take writer round-trip
|
|
|
|
|
const take = {
|
|
|
|
|
name: 'test', frames: 72, width: 320, height: 200, exposure: 2,
|
|
|
|
|
palette: [{ name: 'bg' }, { name: 'skin' }],
|
|
|
|
|
slot: { x: 160, y: 100 },
|
|
|
|
|
parts: [
|
|
|
|
|
{ name: 'head', kind: 'plate', z: 0, interp: 'hold', keys: [{ f: 0, plate: 0 }] },
|
|
|
|
|
{ name: 'mouth', kind: 'poly', z: 30, color: 'skin', interp: 'hold',
|
|
|
|
|
keys: sel.keys.map((k) => ({ f: k.f, src: k.src, pts: shapes[k.src] })) },
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
{ name: 'iris_r', kind: 'disc', z: 22, color: 'iris', interp: 'hold',
|
|
|
|
|
parent: 'eye_r_in', clip: 'eye_r_in',
|
|
|
|
|
keys: [{ f: 0, src: 0, c: { x: 120.4, y: 88.7 }, r: 5.5 }, { f: 2, hidden: true }] },
|
|
|
|
|
{ name: 'pupil_r', kind: 'rect', z: 23, color: 'pupil', interp: 'hold',
|
|
|
|
|
parent: 'iris_r', clip: 'iris_r',
|
|
|
|
|
keys: [{ f: 0, src: 0, c: { x: 120, y: 89 }, size: 3 }] },
|
2026-09-24 14:38:07 -04:00
|
|
|
],
|
|
|
|
|
};
|
|
|
|
|
const text = writeTake(take);
|
|
|
|
|
const keyLines = text.split('\n').filter((l) => l.startsWith('key') && l.includes('n='));
|
|
|
|
|
ok('every key line declares n= matching its point count',
|
|
|
|
|
keyLines.every((l) => {
|
|
|
|
|
const n = +l.match(/n=(\d+)/)[1];
|
|
|
|
|
const pts = l.split(/n=\d+\s+/)[1].trim().split(/\s+/);
|
|
|
|
|
return pts.length === n;
|
|
|
|
|
}), `${keyLines.length} key lines`);
|
|
|
|
|
ok('take declares a plate and a part table',
|
|
|
|
|
/^plate\s+0/m.test(text) && /^part\s+mouth/m.test(text));
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
ok('a disc part declares its clip', /^part\s+iris_r.*clip=eye_r_in/m.test(text));
|
|
|
|
|
ok('a disc key is three integers', /^key\s+iris_r\s+f=0\s+src=0\s+disc=120,89,6$/m.test(text),
|
|
|
|
|
(text.split('\n').find((l) => l.startsWith('key iris_r')) || '').trim());
|
|
|
|
|
ok('a hidden disc key emits hidden', /^key\s+iris_r\s+f=2\s+hidden$/m.test(text));
|
|
|
|
|
ok('a pupil key is a square, not a tessellated polygon',
|
|
|
|
|
/^key\s+pupil_r\s+f=0\s+src=0\s+rect=120,89,3$/m.test(text) &&
|
|
|
|
|
/^part\s+pupil_r.*clip=iris_r/m.test(text));
|
2026-09-24 14:38:07 -04:00
|
|
|
ok('coordinates are integers', !/-?\d+\.\d/.test(text.split('\n').filter((l) => l.startsWith('key')).join('')));
|
|
|
|
|
|
|
|
|
|
return results;
|
|
|
|
|
}
|