diff --git a/README.md b/README.md index 9f1503f..8f13ce7 100644 --- a/README.md +++ b/README.md @@ -17,10 +17,15 @@ modern conveniences belong in the workflow, not the output. See ## Run ```sh -python3 -m http.server 8777 # from this directory -# open http://127.0.0.1:8777 +python3 serve.py # from this directory, then open 127.0.0.1:8777 ``` +Use `serve.py`, not `python3 -m http.server`. The latter sends `Last-Modified` +and browsers cache ES modules on it hard enough that a reload serves a stale +`js/app.js` against a fresh `index.html` — new knobs appear in the markup, nothing +wires them, no error is raised, and the symptom reads as "the feature does not +work". `serve.py` is the same server with `no-store`. + Static files and ES modules — no build step, no dependencies beyond MediaPipe's wasm, which is fetched from a CDN on first use. @@ -43,6 +48,13 @@ playback speed — neither gives a deterministic per-frame pass. assuming, because a guessed fps desynchronises audio from picture — and sync is the one thing this view exists to show. +**Exposure** decides how often the picture gets a new drawing: rip at 24 and +render `on 2s` for 12, `on 3s` for 8. The dense track and the audio are +untouched, so it is a dropdown rather than a re-rip, and the export emits keys +only on the grid instead of the same pose twice. Everything rides the same grid +— mouth, eyes, teeth, plate — because a head cutting on the odd frames while the +mouth cuts on the even ones reads as two performances laid over each other. + **Audio is the playback clock**: `frame = floor(audio.currentTime * fps)`. A slow render loop therefore drops frames instead of drifting, and ½x / ¼x work by setting `playbackRate` with the picture following for free. @@ -62,6 +74,7 @@ is hand-drawn head plates, which this tool does not yet do. | Knob | What it does | | --- | --- | | vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. | +| exposure | How often the picture changes: on 1s, 2s, 3s, 4s. Rip dense, choose timing here. | | mouth lead ±f | Shifts the performance tracks earlier against the audio and the head. `[` `]`. | | contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. | | anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform*. | @@ -99,6 +112,141 @@ kept pixels green, extracted contour amber. Tune against that, not the numbers. brightness and biased low rather than high. Not implemented: it is not visible in the test footage, which reads as a dark cavity with a bright upper-teeth band. +## Eyes + +Three parts per eye, stacked the way the mouth is: a dark **lash ring**, the +**sclera** inside it, and the **iris** inside that, with a square **pupil** in +the iris. The dark ring outside a pale interior is what makes a flat shape read +as an opening rather than a blob, and it is why a blink costs nothing — when the +lid shuts, the traced ring goes flat and the lash line collapses to a lens, +which is a closed eye, drawn correctly, for free. + +The **lids are a feature**, rotoscoped like the mouth: head-local, a key on every +frame, the same `contour avg` knob. They track the face, because the face is what +they are attached to. + +The **iris is a primitive** — a disc at a quantised position — and that is where +the stylisation is. + +### Line of sight + +Gaze is the iris centre relative to the **midpoint of the eye's two corners**, +in units of corner distance. Both corners are in `RIGID`, which is the point: +the origin and the scale are built only from landmarks that do not move under +performance. Measure against the lid ring's centroid instead and every blink +drags that centroid down and fakes a glance at the floor, on exactly the frames +where the eye is most conspicuous. + +**Both eyes share one gaze.** At 320×200 an iris is a handful of pixels and its +centre comes from five landmarks on an eye twenty pixels wide, so the difference +between the two measurements is noise, not vergence — and independent per-eye +noise reads as wall-eyed immediately, which is the most expensive artefact on a +face. Openness stays per-eye, so a wink survives. + +Then the gaze is **quantised to a pixel grid with a dwell**, which is not a +stylisation imposed on the truth: real eyes move in saccades, holding a fixation +and then jumping. The smooth drift left in the measurement is tracker noise plus +head-compensation error, so snapping to a grid and requiring a dwell removes the +noise and recovers the saccade in one operation. The readout reports how many +distinct cells the iris ever occupies — three or four is a character who looks at +things, forty is an unquantised iris sliding around. + +The iris is drawn at the socket read back off the **already-smoothed, already- +subsampled lid ring** — slots 0 and 8 of a 16-slot ring are the corners, and +subsampling to any even budget keeps them at output indices 0 and `n/2`. So the +iris is placed in the frame of the exact polygon it sits inside and cannot drift +relative to its own eye. Size is authored from the take's mean eye width, not +remeasured per frame: a radius that breathes by a fraction of a pixel flickers a +pixel on and off around the whole silhouette. + +The iris is **stencilled to the sclera** and the pupil to the iris — the indexed +buffer is its own clip mask, the way Animator Pro would do it. So the lid crops +the iris at extreme gaze automatically, and nothing needs to clamp the gaze, +which would flatten the performance at exactly the extremes that carry it. + +### Blinking + +Openness is the lid gap over the corner distance — normalised, so one threshold +carries across takes and faces. It gets hysteresis and a dwell like the teeth, +plus one knob the teeth do not have: **blink hold**. A blink is 100–150ms, which +is one frame at 12fps, and a single frame of closed eye reads as a dropped frame +rather than as a blink. Animators draw a blink over two or three drawings for +that reason, so once the eye shuts it stays shut for `hold` frames. + +### The pupil is a square + +At this size a pupil is three pixels across, and a circle of radius 1.5 is not a +circle — it is a plus sign with the corners gnawed off, and it changes shape as +it moves. A square that size is a deliberate mark that stays the same mark +wherever it lands. It is drawn from a rounded centre shared with the iris, so it +is exactly its nominal size on every frame instead of spilling to the next pixel +on some and not others. + +### Which iris is which + +The refined mesh appends ten iris points, five per eye, and MediaPipe's own +left/right naming is viewer-relative in some places and subject-relative in +others. Getting it backwards swaps the irises, which looks *almost* right — each +eye still has a disc roughly where it belongs — so it survives an eyeball and +then reads as a subtly wall-eyed character forever. The pairing is therefore +**resolved from the geometry**, by voting each block's distance to each eye's +corner midpoint across every frame, and the selftest feeds it a track built the +other way round to prove it actually looks. + +| Knob | What it does | +| --- | --- | +| eye vertices | Lid ring vertex budget, off a 16-slot ring. | +| lash line | How far the dark ring sits outside the lid, in pixels. | +| blink cut | Openness below which the eye is shut. Normalised by corner distance. | +| blink hold | Minimum frames a blink stays on screen. A one-frame blink is a dropout. | +| blink dwell | Frames a change must persist. Usually 0 — unlike the teeth, a real blink *is* one frame. | +| gaze gain | Exaggerates or damps the throw. Measured excursion is small; a character usually wants more. | +| gaze step | The pixel grid the iris snaps to. 0 = off, and then dwell does nothing either. | +| gaze dwell | How long a new cell must hold before it takes. Together with step, this is what makes saccades. | +| iris size | Diameter as a percentage of eye width. | +| pupil | Square pupil in whole pixels. 0 = off. | + +## Brows + +A brow at 320×200 is about fourteen pixels wide and three tall. Its **shape** +carries almost nothing at that size; its **height above the eye** carries the +expression, and a brow raise is the most legible beat on a face. So the ring is +traced and the height is quantised — the same split the eyes got, where the lid +is a traced feature and the iris a quantised primitive. + +The decomposition matters. The traced ring already contains the real height, so +adding a quantised raise on top would move the brow twice. Instead the height is +measured *out* of the ring, quantised, and put back: the shape that renders is +his, at a height that snaps between a few levels and holds. + +Height is measured at **both ends**, not as one number, because raise and tilt +are different expressions out of one mechanism — both ends up is surprise, inner +up alone is worry, inner down is anger. They share a dwell, so the brow hits its +pose in one frame instead of crawling into it with one end arriving first. + +It is measured against the eye's **corner midpoint**, never its lid — the same +trap the gaze origin has, and worth avoiding twice: brows and lids move together +constantly, so a brow that jumped on every blink would read as a tic. The rest +pose comes from the take **median**, not the neutral frame, for the same reason +gaze does: that frame is chosen by minimum mouth aperture and says nothing +whatever about the brows. + +Two correspondences are resolved from geometry rather than declared: which ring +is which brow, and which end of a ring is the outer one. The second matters more +— get it backwards and the tilt mirrors, so worry renders as its own opposite, +which reads as a directed performance choice and would never be questioned. +Which *edge* of the brow is the upper one is deliberately left unresolved: it +traverses the same ring the other way round, an even-odd fill has no winding, +and the two ends still land on fixed slots either way. + +| Knob | What it does | +| --- | --- | +| brow vertices | Ring vertex budget, off a 10-slot ring. | +| brow weight | Thickens the ring outward. It needs it at three pixels tall. | +| brow raise gain | Exaggerates or damps the raise. | +| brow step | The pixel grid the height snaps to. 0 = off. | +| brow dwell | How long a new height must hold. Shared across both ends. | + ## The plate is reference, not art The plate layer has several representations because its job changes. Cycle with @@ -124,6 +272,45 @@ egg by construction, and no landmark precision fixes that. Hence the photo. **Save frame 4x** writes the current registered composite as a 1280×800 PNG to draw on. +## Paint — background cels + +**A sketch.** It exists to test whether the aesthetic holds when a human draws +the background instead of the tracker deriving it, and it is meant to be +replaced by a real paint surface with onion skin and undo. It is one dependency- +free module, `js/paint.js`, so throwing it away is a delete rather than surgery. + +Cels are drawn on the frames that get their own drawing and **hold until the +next one** — the same rule the plate follows, and literally the same lookup. You +can scrub anywhere and keep drawing on the cel you can see; the header says +which one you are editing and how far it holds. + +- **pen** — click to place vertices, click the green box on the first one (or + Enter / double-click) to close. +- **edit** — click a shape to select, drag a vertex or the whole shape, + Shift-click an edge to insert a vertex, Alt-click one to + remove it, Del to delete the layer. +- **Layers** stack Photoshop-style, front at the top, with per-layer colour, + show/hide and reorder. +- **Copy previous** brings the last drawing forward onto this frame. It means + the nearest earlier enabled frame that *actually has* a drawing, skipping the + empty ones — every frame is enabled until you thin the strip out, so the naive + rule resolved to `f-1` and it looked like it only ever copied the frame to the + left. **Drag a frame** from the strip onto the canvas to seed from any other + frame instead. Both deep-copy; the two cels never share point arrays. +- Frames carrying a drawing are marked **▣** in the strip, so you can see the + rhythm rather than having to remember it. + +Two rules are enforced rather than left to discipline. Colours are **palette +indices**, so you cannot pick one that is not in the ramp — sampling colour from +the source is the one move `docs/design.md` says is irrecoverable. And vertices +**snap to the 320×200 grid**, because on a hard-edged indexed rasteriser a shape +nudged by 0.4px moves an edge by a whole pixel or not at all depending on where +it lands, which shimmers instead of holding. + +Drawings autosave to `localStorage` per take name. They are the only thing in +the tool a person made by hand; everything else regenerates. They are not in the +`.take` export yet. + ## Two kinds of sparseness Sparseness has two unrelated causes, and conflating them was the original design @@ -131,6 +318,10 @@ error here. **Aesthetic** sparseness is set by the extraction rate — pick 12fp you have already chosen your timing. **Labour** sparseness is a human drawing each one, and it binds only on the plate. +Aesthetic sparseness is the **exposure** control, not the extraction rate — +making it a render-time grid means auditioning 12 against 24 costs a dropdown +instead of a re-rip and a full re-detection. + So the mouth keeps **every** frame: it is traced, and therefore free. In limited animation lip sync is routinely the densest element, on 1s, while heads hold on 2s and 3s. @@ -162,7 +353,7 @@ chromium --headless --virtual-time-budget=8000 --dump-dom \ http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+' ``` -Or open `selftest.html`. 41 assertions over the stages below detection, plus a +Or open `selftest.html`. 105 assertions over the stages below detection, plus a wiring cross-check: every `el('id')` in `app.js` must exist in `index.html`. A knob wired in one but not the other throws during wiring, which aborts the rest of the module and leaves a blank page — a symptom that points nowhere near its @@ -176,7 +367,7 @@ eyeball. ## Not done yet -Eyes and irises; hand-drawn head plates and per-plate mouth slots (the strip +Hand-drawn head plates and per-plate mouth slots (the strip decides *which frames need one*, but you cannot yet supply the drawing); real performer→character calibration (currently identity, fitting the face oval to the canvas); the override layer; anything on the Animator Pro side. The plate is a diff --git a/audio.wav b/audio.wav index 9c39a31..f5793b0 100644 Binary files a/audio.wav and b/audio.wav differ diff --git a/docs/design.md b/docs/design.md index a614389..9fd57d9 100644 --- a/docs/design.md +++ b/docs/design.md @@ -52,9 +52,17 @@ artist or a fixed authored table. | Kind | Source | Vocabulary | Interp | | --- | --- | --- | --- | | **Plate** — head, hair, body | Hand-drawn | Closed: a few drawings per character | hold | -| **Feature** — mouth, lids | Rotoscoped from landmarks | Open: derived from this take | hold | +| **Feature** — mouth, lids, brows | Rotoscoped from landmarks | Open: derived from this take | hold | | **Interior** — mouth interior, teeth | Image content within a feature | Open | hold | | **Primitive** — iris | Landmark centroid as a disc | Quantised | hold | +| **Scalar** — brow raise, gaze | One number out of a feature | Quantised | hold | + +The last row took the longest to see. A brow is a feature *and* a scalar: the +ring is traced because the shape should be his, but at three pixels tall the +shape carries almost nothing while the height above the eye carries the +expression. So the height is measured out of the traced ring, quantised, and put +back. Extracting the scalar without removing it first would move the part twice, +because the traced ring already contains the height. The asymmetry is deliberate, and it is the opposite choice in each case. @@ -70,9 +78,19 @@ avoidance. ## Two kinds of sparseness -Conflating these was the original design error. **Aesthetic** sparseness is set -by the extraction rate: pick 12fps and the timing is already chosen. **Labour** -sparseness is a human drawing each one, and it binds only on the plate. +Conflating these was the original design error. **Aesthetic** sparseness is the +rate the picture changes at. **Labour** sparseness is a human drawing each one, +and it binds only on the plate. + +Aesthetic sparseness used to be set by the extraction rate — rip at 12 and the +timing is chosen. That was wrong in a small way: it makes the timing a property +of a directory of PNGs, so auditioning 12 against 24 means re-ripping the clip +and re-running detection over the whole of it, and the decision you most want to +play with is the one that costs the most to change. Rip at the camera's rate and +quantise at render time instead — an **exposure** grid, on 1s, 2s, 3s — so the +dense track keeps everything, the audio clock is untouched, and the timing is a +dropdown rather than a re-rip. The take format already carried an `exposure` +field for this; it was simply never driven. So the mouth keeps **every** frame — it is traced, and therefore free. In limited animation lip sync is routinely the densest element, on 1s, while heads hold on @@ -136,6 +154,83 @@ Two escapes, both used: temporal smoothing is well defined, and the star-shaped result suits flat colour. +## Every part is measured in the frame of the thing it is attached to + +The mouth is expressed against the head. The iris is expressed against its own +eye — specifically against the midpoint of that eye's two corners, in units of +corner distance. Both corners are rigid landmarks, so the origin and the scale +of the measurement are immune to the performance being measured. Against the lid +ring's centroid instead, every blink would drag the origin down and fake a glance +at the floor on exactly the frames where the eye is most visible. + +The rule generalises: **measure a feature in a frame built only from landmarks +that do not move with it.** It is the same argument as "rigid landmarks only" for +the anchor fit, one level down. + +There is a tempting over-application. An eye can be pinned into a fixed socket +fitted to its corners' mean over the shot, which removes the residual wobble a +2D similarity cannot — and it is wrong. That residual is real motion of the eye +relative to the head, it is still there in the footage, and removing it leaves +the drawn eyes hanging still over a registered photo whose eyes are moving. A +part must track the face in the same space the underlay is drawn in. The wobble +is a job for the bounded contour average below, not for a second anchor. + +Placement follows from the same idea. The iris is drawn in the frame of the +already-smoothed, already-subsampled lid ring, read off the ring's own corner +vertices, so it cannot drift relative to the eye it sits inside and it inherits +the contour average for free. Size, by contrast, is authored from the take's +mean, never remeasured per frame: a radius that breathes by a fraction of a pixel +flickers a pixel on and off around the whole silhouette. + +## The indexed buffer is its own stencil + +Parts that nest — iris inside sclera, pupil inside iris — clip by colour key: +paint only where the buffer already holds the parent's index. This is how +Animator Pro would do it, it costs one comparison per pixel, and it composes +transitively, so a blink takes the right bite out of the pupil without anything +computing where. + +It also removes a temptation. Without a stencil the gaze has to be clamped to +keep the iris inside the lid, and a clamp flattens the performance at exactly the +extremes that carry it. + +## Quantisation can be the truthful choice + +Gaze snapped to a pixel grid with a dwell is the "Primitive — quantised" row of +the part table, and it looks like a stylisation imposed on a continuous +measurement. It is not. Real eyes move in saccades: hold a fixation, jump, hold. +The smooth drift left in the measured signal is tracker noise plus +head-compensation error. Snapping to a grid and requiring a dwell removes the +noise and recovers the saccade in the same operation — the rare case where the +aesthetic rule and the physiology agree. + +The count of distinct cells the iris ever occupies is the number the knobs exist +to control. Three or four is a character who looks at things; forty is an +unquantised iris sliding around. + +## Some thresholds need a minimum duration, not just a dwell + +A dwell delays a change until it has persisted, which is the right guard against +chatter and is what the teeth use. A blink needs the opposite guard as well. It +lasts 100–150ms — one frame at 12fps — and a single frame of closed eye reads as +a dropped frame rather than as a blink. Animators draw a blink over two or three +drawings for that reason, so once the eye shuts it must stay shut for a minimum +number of frames. Detection accuracy is not the problem; legibility is. + +## Resolve correspondences from data when a wrong guess is survivable + +The refined mesh appends ten iris points, five per eye, and the upstream +left/right naming is viewer-relative in some documentation and subject-relative +in others. Swapping them looks *almost* right — each eye still has a disc roughly +where it belongs — so the error survives inspection and then reads as a subtly +wall-eyed character for the life of the project. + +A hardcoded table is the wrong shape for a fact like that. Voting each block's +distance to each eye's corner midpoint across every frame settles it from the +geometry, cannot be got wrong, and keeps working if the model is renumbered. The +test feeds it a track built the other way round, because a resolver checked only +against the convention it was written for is checking nothing. + ## The bounded smoothing exception *Smooth the transform, never the contour* held while keys were sparse: sampling @@ -188,6 +283,7 @@ Current modules: | Module | Role | | --- | --- | | `landmarks.js` | Index tables. Ring arrays are ordered traversals: slot position *is* vertex identity. | +| `pipeline.js` | …also eye openness, gaze, blink resolution and the iris pairing vote. | | `mathutil.js` | Similarity fit, Procrustes mean, temporal smoothing. | | `pipeline.js` | Stabilise → subsample → key-select → frame-removal. | | `interior.js` | Teeth from image content: Otsu, morphology, components, radial contour. | @@ -202,7 +298,7 @@ Current modules: - **A paint surface.** The plates have nowhere to be drawn. This is the largest gap between "tool" and "suite": a pixel paint canvas with onion skin, palette constraint, and the registered underlay behind it. -- Eyes, irises, brows as parts. +- A tongue. - Plate libraries with per-plate mouth slots. - Real performer→character calibration (currently identity). - The override layer. diff --git a/index.html b/index.html index cb59dd6..31b4e0c 100644 --- a/index.html +++ b/index.html @@ -42,6 +42,8 @@ .fr.drop { border-color:#2a2f3e; } .fr.drop canvas { opacity:.26; filter:grayscale(1); } .fr.cur { border-color:var(--accent); } + .fr.cel::before { content:'▣'; position:absolute; left:3px; top:1px; color:#c084fc; + font-size:10px; text-shadow:0 0 2px #000; } .fr.keep::after { content:'●'; position:absolute; right:3px; top:1px; color:var(--ok); font-size:10px; } #sheet { display:flex; flex-wrap:wrap; gap:10px; } @@ -52,6 +54,19 @@ .sw { display:flex; align-items:center; gap:5px; color:var(--dim); font-size:11px; } .sw input { width:26px; height:20px; padding:0; border:1px solid var(--line); background:none; } .legend { color:var(--dim); font-size:11px; margin-top:6px; } + /* paint: a sketch, see js/paint.js */ + #cv-paint { border:1px solid var(--line); cursor:crosshair; touch-action:none; } + #cv-paint.drop { border-color:var(--ok); } + #paintlayers { flex:0 0 280px; max-height:600px; overflow:auto; } + .lay { display:flex; align-items:center; gap:5px; padding:3px 4px; border-radius:3px; + border:1px solid transparent; cursor:pointer; } + .lay.sel { border-color:var(--accent); background:#1c2130; } + .lay i { width:12px; height:12px; border-radius:2px; border:1px solid #0006; flex:0 0 auto; } + .lay b { flex:1; font-weight:400; color:var(--dim); font-size:11px; } + .lay.sel b { color:var(--fg); } + .lay select { font:inherit; font-size:10px; background:#1c2130; color:var(--dim); + border:1px solid var(--line); border-radius:2px; padding:1px 2px; max-width:88px; } + .lay button { padding:0 5px; font-size:11px; line-height:18px; } kbd { background:#1c2130; border:1px solid var(--line); border-radius:2px; padding:0 4px; color:var(--fg); font-size:11px; } @@ -63,6 +78,21 @@ + + + @@ -86,18 +116,23 @@

source + landmarks

-
outer lip — · inner lip —
+
outer lip — · inner lip — · + lids — · iris — · + brows —

stabilised (head-local)

-
should sit still except the mouth · grey = held plate outline
dark green ghost = unshifted mouth when lead ≠ 0
+
should sit still except the mouth and eyes · grey = held plate outline
+ dark green ghost = unshifted mouth when lead ≠ 0 · red lid ring = blink

flat render — 320×200 indexed

audio drives the clock — dropped frames, never drift
+ exposure holds the picture on a grid: rip at 24, render on 2s for + 12. The dense track and the audio are untouched, so it is reversible
plate representation: B cycles · photo modes are registered into raster space, so tracing them lands on the mouth
@@ -132,6 +167,21 @@ + + + + + + + + + + + + + + +
mouth lead shifts the performance tracks earlier (positive) against @@ -150,10 +200,34 @@ brightness; prefer upper biases component choice toward the top of the cavity, where teeth are and the tongue is not. dwell is how many frames a presence change must persist.
+ blink cut is lid gap over corner distance — normalised, so one + value carries across takes. blink hold is the minimum length of a + blink: a real blink is one frame at 12fps and a single frame of closed + eye reads as a dropout, so it is extended to a beat. + brow step and brow dwell quantise the brow's HEIGHT above + the eye, not its shape — the traced ring is his, the height snaps between + a few levels and holds. Both ends move independently, so raise and tilt + come out of one control: both up is surprise, inner up is worry, inner + down is anger. brow weight thickens the ring, which it needs at + three pixels tall.
+ pupil is a square, in whole pixels, 0 to turn it off: at this size + a circle of radius 1.5 is a plus sign with the corners gnawed off and it + changes shape as it moves, where a square stays the mark you drew.
+ gaze step is the grid the iris snaps to, in raster pixels, and + gaze dwell is how long a new cell must hold — together they turn + drift into saccades. gaze gain exaggerates or damps the throw; + measured excursion is small and a character usually wants more of it.
suggest tolerance only affects the Suggest button: max head movement allowed before a new drawing is required.
+
+

gaze field

+ +
+
green = every cell the iris visits in the take · + grey = raw · amber = where it is now, quantised
+

teeth measurement

@@ -170,6 +244,51 @@
+
+

paint — background cels, drawn on the frames that get their own drawing

+
+ + + + + + +
+
+ +
+
+
+
+
+ pen click to place vertices · click the green box on the first one, or + Enter / double-click, to close · Esc cancels · + Backspace drops the last point
+ edit click a shape to select · drag a vertex or the shape itself · + Shift-click an edge inserts a vertex · Alt-click a vertex + removes it · Del deletes the layer
+ opacity dims the drawing over the registered source frame so you can + trace it. It is an editing aid only — never rasterised, never exported, and + the flat render is untouched. At 100% the view is exactly the normal + composite; below it the placeholder oval is dropped, since it covers the + face you turned the slider down to see.
+ Copy previous takes the last enabled frame that actually has a drawing + — skipping empty ones, so it behaves the same before and after you thin the + strip out. Frames carrying a drawing are marked ▣.
+ drag a frame from the strip onto the canvas to copy its drawing, every + layer. Vertices snap to the 320×200 grid, and colours are palette indices — + you cannot pick one that is not in the ramp. +
+
+

drawings needed — registered reference per kept frame, with the range it holds

diff --git a/js/app.js b/js/app.js index 752bac5..79e9a0e 100644 --- a/js/app.js +++ b/js/app.js @@ -1,12 +1,17 @@ import { FaceLandmarker, FilesetResolver } from 'https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@1.0.1/vision_bundle.mjs'; -import { LIPS_OUTER, LIPS_INNER, FACE_OVAL } from './landmarks.js'; -import { stabilize, toRasterRing, smoothContours, suggestPlateFrames, heldFrame, shiftIndex } from './pipeline.js'; +import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, + EYE_R_RING, EYE_L_RING, IRIS_A, IRIS_B, + BROW_A_RING, BROW_B_RING } from './landmarks.js'; +import { stabilize, toRasterRing, smoothContours, suggestPlateFrames, heldFrame, shiftIndex, + exposeIndex, eyeSignals, gazeOrigin, quantizeSnap, resolveBlink, + browSignals } from './pipeline.js'; import { IndexedRaster } from './raster.js'; import { drawRegistered, posterizeInto } from './underlay.js'; import { extractTeeth } from './interior.js'; -import { applySim } from './mathutil.js'; +import { applySim, offsetRing } from './mathutil.js'; import { writeTake } from './take.js'; import { synthDense } from './synth.js'; +import { PaintUI, drawCel, cloneCel, newCel } from './paint.js'; const RW = 320, RH = 200, ZOOM = 2, THUMB = 92; @@ -16,8 +21,25 @@ const PALETTE = [ { name: 'skin_dark', hex: '#7a4f3a' }, { name: 'mouth_dark', hex: '#24161a' }, { name: 'teeth', hex: '#d9cfc2' }, + // Sclera is not white, and that is authored, not measured. A true white at + // 320x200 next to a warm skin ramp reads as a hole punched in the face; the + // eye sits in a socket, in shadow, so it is a dimmer and cooler tone than the + // teeth, which catch the light. The iris is one dark tone: at this size an + // iris is about five pixels across and a pupil inside it would be one, so the + // iris IS the pupil. Resolving it further would be drawing detail the format + // cannot hold. + { name: 'eye_white', hex: '#c9c3b4' }, + // Three tones for the eye - sclera, iris, pupil - which is the "two or three + // tones per part" budget, spent where it buys the most: an eye with no tonal + // step inside it reads as a hole. + { name: 'iris', hex: '#4a5468' }, + { name: 'pupil', hex: '#171a22' }, + // Brows get their own entry rather than sharing skin_dark with the lash line. + // They are hair, not shadow: when hair plates exist they want to match those, + // and tying them to the lash means you cannot change one without the other. + { name: 'brow', hex: '#3a2a22' }, ]; -const IDX = { bg: 0, base: 1, dark: 2, mouth: 3, teeth: 4 }; +const IDX = { bg: 0, base: 1, dark: 2, mouth: 3, teeth: 4, white: 5, iris: 6, pupil: 7, brow: 8 }; const state = { dense: null, images: [], stab: null, xform: null, @@ -27,8 +49,13 @@ const state = { fps: 12, audio: null, // fps comes from manifest.json, never guessed aspect: 1, // imgW/imgH; converts MediaPipe's anisotropic space lead: 0, // performance-track offset in frames + exposure: 1, // 1 = on 1s, 2 = on 2s. Picture holds; audio does not. interior: null, // per-frame teeth measurement from image content teeth: null, // resolved per-frame {show, t} after knobs + eyes: null, // resolved per-frame lid rings, shut flags, iris discs + brows: null, // resolved per-frame brow rings after quantised raise + cels: new Map(), // kept frame -> hand-painted background layers + eyeSig: null, // raw eye measurement, kept for the gaze readout }; const el = (id) => { @@ -53,6 +80,24 @@ const opts = () => ({ contourSmooth: +el('contourSmooth').value, apertureThresh: +el('apertureThresh').value / 1000, tol: +el('tol').value / 1000, + exposure: +el('exposure').value, + browVerts: +el('browVerts').value, + browWeight: +el('browWeight').value, + browGain: +el('browGain').value / 100, + browStep: +el('browStep').value, + browDwell: +el('browDwell').value, + irisAnchor: el('irisAnchor').value, + gazeOrigin: el('gazeOrigin').value, + eyeVerts: +el('eyeVerts').value, + lashPx: +el('lashPx').value, + irisSize: +el('irisSize').value / 100, + gazeGain: +el('gazeGain').value / 100, + gazeStep: +el('gazeStep').value, // whole raster pixels + pupilPx: +el('pupilPx').value, + gazeDwell: +el('gazeDwell').value, + blinkCut: +el('blinkCut').value / 1000, + blinkHold: +el('blinkHold').value, + blinkDwell: +el('blinkDwell').value, }); function status(msg, kind = '') { @@ -167,6 +212,7 @@ function rebuild(resetKeep) { if (!state.dense) return; const o = opts(); state.lead = o.lead; + state.exposure = o.exposure; const N = state.dense.length; state.stab = stabilize(state.dense, o.smoothWin, state.aspect); @@ -193,6 +239,8 @@ function rebuild(resetKeep) { state.extractKey = extractKey(o); } state.teeth = resolveTeeth(o); + state.eyes = buildEyes(o); + state.brows = buildBrows(o); // Plate outline per frame, so a kept frame shows its own head shape. state.plates = state.stab.oval.map((r) => r.map(state.xform)); @@ -220,6 +268,188 @@ function faceBoxes() { }); } +// Eyes: lid rings traced per frame, blinks resolved per eye, one gaze shared. +// +// Lids are a FEATURE in the part table - rotoscoped, open vocabulary, a key on +// every frame - so they get exactly the mouth's treatment, including the same +// bounded contour average. The iris is a PRIMITIVE: a disc whose position is +// quantised, which is where the stylisation lives. +function buildEyes(o) { + const st = state.stab, N = state.dense.length; + const sig = eyeSignals(st); + state.eyeSig = sig; + + const blink = { cut: o.blinkCut, dwell: o.blinkDwell, hold: o.blinkHold }; + const shutR = resolveBlink(sig.openR, blink); + const shutL = resolveBlink(sig.openL, blink); + + // Head-local, subsampled, contour-averaged - the identical chain the mouth + // takes, with the identical knob. The eye tracks the face, because the face + // is what it is attached to; what gets removed is per-frame detector jitter, + // not the motion. + const lidR = smoothContours( + st.lidR.map((r) => toRasterRing(r, EYE_R_RING, o.eyeVerts, state.xform)), o.contourSmooth); + const lidL = smoothContours( + st.lidL.map((r) => toRasterRing(r, EYE_L_RING, o.eyeVerts, state.xform)), o.contourSmooth); + + // Where the iris hangs. Three behaviours, because this turns out to be an + // aesthetic choice and not only a correctness one. + // + // STEADY (default) reads the socket back off the DRAWN ring. Slots 0 and 8 of + // a 16-slot lid ring are the two corners, and subsampling to any even budget n + // keeps them at output indices 0 and n/2 - so the ring that gets rendered + // carries its own corners with it. The iris is then placed in the frame of the + // exact polygon it sits inside, after smoothing, after subsampling: it cannot + // drift relative to its own eye, and it inherits the contour average for free. + // + // FREE reads the raw per-frame corners instead, jitter and all. It is what the + // eyes did before any of this, and it is not simply worse - the detector noise + // reads as liveliness, the eye never sits perfectly still, and against flat + // hand-drawn plates that restlessness can be the thing that sells it. It is + // also the honest baseline to compare the other two against. + // + // LOCKED pins the socket to the take's mean, so the eye never moves in the + // head at all. Watch it against a photo underlay and the drawn eyes hang still + // over a face whose eyes are moving - that is the registration cost, and it is + // real - but once the plate is a drawing rather than a photograph, nothing is + // being registered against and it reads as a deliberately locked-off stare. + const ringSocket = (ring) => { + const a = ring[0], b = ring[ring.length / 2]; + return { cx: (a.x + b.x) / 2, cy: (a.y + b.y) / 2, w: Math.hypot(a.x - b.x, a.y - b.y) }; + }; + const rawSocket = (corners, f) => { + const a = state.xform(corners[f][0]), b = state.xform(corners[f][1]); + return { cx: (a.x + b.x) / 2, cy: (a.y + b.y) / 2, w: Math.hypot(a.x - b.x, a.y - b.y) }; + }; + const meanSocket = (rings) => { + const acc = rings.reduce((a, r) => { + const k = ringSocket(r); + return { cx: a.cx + k.cx, cy: a.cy + k.cy, w: a.w + k.w }; + }, { cx: 0, cy: 0, w: 0 }); + const n = rings.length; + return { cx: acc.cx / n, cy: acc.cy / n, w: acc.w / n }; + }; + const socketFor = (rings, corners) => { + if (o.irisAnchor === 'locked') { const k = meanSocket(rings); return () => k; } + if (o.irisAnchor === 'free') return (f) => rawSocket(corners, f); + return (f) => ringSocket(rings[f]); + }; + const skR = socketFor(lidR, st.cornersR), skL = socketFor(lidL, st.cornersL); + const socket = (ring) => ringSocket(ring); + + // Iris radius comes from the take's MEAN eye width, not the current frame's. + // Size is authored; only position is tracked. A radius recomputed per frame + // would breathe by a fraction of a pixel as the fit's depth-scale wanders, + // and at this resolution a fraction of a pixel is a pixel flicking on and off + // around the whole silhouette. + const meanW = (rings) => rings.reduce((a, r) => a + ringSocket(r).w, 0) / rings.length; + const wR = meanW(lidR), wL = meanW(lidL), w = (wR + wL) / 2; + + // Calibrate against the neutral, apply the artist's gain, and only then + // quantise - the grid should be a grid of DRAWN positions, because that is + // what a viewer reads. Gain is an authored parameter: measured gaze excursion + // is small and a character's eye usually wants more throw than a performer's, + // which is a decision for a person and not for the detector. + const origin = gazeOrigin(sig.gazeRaw, o.gazeOrigin, state.neutral); + state.gazeOriginValue = origin; + const px = sig.gazeRaw.map((g) => ({ + x: (g.x - origin.x) * o.gazeGain * w, + y: (g.y - origin.y) * o.gazeGain * w, + })); + const gaze = quantizeSnap(px, o.gazeStep, o.gazeDwell); + + const eye = (sk, lids, shut, rad, f) => { + const e = sk(f); + return { + // The lash line is the lid ring pushed outward by a fixed number of + // pixels, exactly as the mouth's outer ring sits outside its inner one. + // When the eye shuts, the traced ring goes near-degenerate and this + // collapses to a lens - which is a closed eye, drawn correctly, for free. + lash: offsetRing(lids[f], o.lashPx), + lid: lids[f], + shut: shut[f], + // Rounded to whole pixels. The rasteriser quantises everything anyway, so + // this costs nothing - but it means the iris and the square pupil share + // one integer centre, so the pupil is exactly its nominal size on every + // frame instead of spilling to the next pixel on some and not others. + iris: { x: Math.round(e.cx + gaze[f].x), y: Math.round(e.cy + gaze[f].y), r: rad }, + pupil: o.pupilPx, + }; + }; + + return { + gazePx: px, gaze, shutR, shutL, hasIris: sig.hasIris, + frames: Array.from({ length: N }, (_, f) => ({ + r: eye(skR, lidR, shutR, (wR * o.irisSize) / 2, f), + l: eye(skL, lidL, shutL, (wL * o.irisSize) / 2, f), + })), + }; +} + +// Brows: ring traced every frame, HEIGHT quantised. +// +// The decomposition is the point. The traced ring already contains the brow's +// real height, so adding a quantised raise on top would move it twice. Instead +// the height is measured out of the ring, quantised, and put back - the shape +// that renders is his, at a height that snaps between a few authored levels and +// holds. That is the same split the eyes got: lid traced as a feature, iris +// position quantised as a primitive. +// +// Two ends, not one height, warped linearly between them. Raise and tilt are +// different expressions out of one mechanism: both ends up is surprise, inner +// up alone is worry, inner down is anger. +function buildBrows(o) { + const st = state.stab, N = state.dense.length; + const sig = browSignals(st); + state.browSig = sig; + + const ringOf = (side) => (side === 'R' ? sig.pairing.right : sig.pairing.left); + const table = (side) => (ringOf(side) === 'browA' ? BROW_A_RING : BROW_B_RING); + + const build = (side, corners) => { + const rings = smoothContours( + st[ringOf(side)].map((r) => toRasterRing(r, table(side), o.browVerts, state.xform)), + o.contourSmooth); + + // Eye width in raster pixels, so the raise converts from eye widths into the + // units the grid is expressed in and the knob means the same on any framing. + const wpx = (f) => { + const a = state.xform(corners[f][0]), b = state.xform(corners[f][1]); + return Math.hypot(a.x - b.x, a.y - b.y); + }; + const meanW = st.cornersR.reduce((a, _, f) => a + wpx(f), 0) / N; + + // Rest pose from the take MEDIAN, never from the neutral frame. That frame + // is chosen by minimum mouth aperture and says nothing about the brows, and + // the same mistake on the gaze origin re-pointed an entire performance. + const rest = gazeOrigin(sig[side], 'median'); + const px = sig[side].map((g) => ({ + x: (g.x - rest.x) * o.browGain * meanW, + y: (g.y - rest.y) * o.browGain * meanW, + })); + const q = quantizeSnap(px, o.browStep, o.browDwell); + + const frames = rings.map((ring, f) => { + // Raise is measured upward but y grows downward, so a positive raise is a + // negative y offset. + const dOuter = -(q[f].x - px[f].x), dInner = -(q[f].y - px[f].y); + const a = state.xform(corners[f][0]), b = state.xform(corners[f][1]); + const span = b.x - a.x; + const warped = ring.map((p) => { + // Position along the brow's own axis, outer end to inner end. Taken from + // x against the eye corners rather than from ring slots, because + // subsampling does not keep the end slots at any given budget. + const t = span === 0 ? 0 : Math.min(1, Math.max(0, (p.x - a.x) / span)); + return { x: p.x, y: p.y + dOuter + (dInner - dOuter) * t }; + }); + return offsetRing(warped, o.browWeight); + }); + return { frames, px, q }; + }; + + return { R: build('R', st.cornersR), L: build('L', st.cornersL), pairing: sig.pairing }; +} + // Presence gets hysteresis and a minimum dwell, the same treatment plate // selection gets: a teeth block that blinks on and off for single frames is // worse than one that is simply absent. Appearing needs a clear signal, staying @@ -275,14 +505,34 @@ function resolveTeeth(o) { // stand-in until a drawing exists. function renderFrame(f, mode = plateMode()) { const r = new IndexedRaster(RW, RH); - const pf = heldFrame(keptSorted(), f); // the plate frame on screen + const pf = plateIndex(f); // the plate frame on screen if (mode === 'posterize' && state.images[pf]) { posterizeInto(r, state.images[pf], state.stab.transforms[pf], state.xform, PALETTE.map((p) => p.hex)); } else { r.clear(IDX.bg); - if (mode === 'oval' || mode === 'oval+photo') r.fillPoly(state.plates[pf], IDX.base); + } + // Painted cels sit BEHIND the face and hold on the same frames the plate + // does - pf is already "the most recent kept frame at or before f", which is + // exactly the rule the user draws against: a cel holds until the next frame + // that has its own drawing. + drawCel(r, state.cels.get(pf), (i) => i); + if (mode !== 'posterize' && (mode === 'oval' || mode === 'oval+photo')) { + r.fillPoly(state.plates[pf], IDX.base); + } + + // Eyes run on the CLOCK, not on the mouth lead. The lead is a lip-sync + // device: it exists because a mouth shape anticipates the sound it makes. + // Nothing about a blink or a glance is tied to the audio, so shifting the + // eyes would only slide them off the head that carries them. + // Eyes and brows ride the exposure grid but NOT the mouth lead: the lead is a + // lip-sync device and nothing about a blink or a brow is tied to the audio. + const ef = perfIndex(f); + if (state.eyes) drawEyes(r, state.eyes.frames[ef]); + if (state.brows) { + r.fillPoly(state.brows.R.frames[ef], IDX.brow); + r.fillPoly(state.brows.L.frames[ef], IDX.brow); } const mf = leadIndex(f); // performance frame, possibly ahead @@ -295,13 +545,34 @@ function renderFrame(f, mode = plateMode()) { return r; } +// Lash ring, then sclera, then iris - the same three-layer structure the mouth +// has, for the same reason: the dark ring outside the pale interior is what +// makes a flat shape read as an opening rather than a blob. +// +// The iris is stencilled to the sclera it was just drawn over, so the lid crops +// it automatically. Nothing needs to clamp the gaze to keep the iris inside the +// eye, which matters because a clamp would flatten the performance at exactly +// the extremes that carry it. +function drawEyes(r, e) { + for (const s of [e.r, e.l]) { + r.fillPoly(s.lash, IDX.dark); + if (s.shut) continue; // a shut eye IS the lash line, alone + r.fillPoly(s.lid, IDX.white); + r.fillDisc(s.iris.x, s.iris.y, s.iris.r, IDX.iris, IDX.white); + // Stencilled to the iris, which is itself stencilled to the sclera - so the + // pupil is cropped by the lid transitively, and a blink or an extreme gaze + // takes the right bite out of it without anything having to compute where. + if (s.pupil) r.fillRect(s.iris.x, s.iris.y, s.pupil, IDX.pupil, IDX.iris); + } +} + const plateMode = () => el('plateMode').value; // Photo modes composite under the indexed layer, so the flat shapes stay exactly // as they render while the reference sits behind them. function compositeRender(canvas, f, zoom) { const mode = plateMode(); - const pf = heldFrame(keptSorted(), f); + const pf = plateIndex(f); const img = state.images[pf]; const showPhoto = img && (mode === 'photo' || mode === 'photo-dim' || mode === 'oval+photo'); @@ -315,23 +586,31 @@ function compositeRender(canvas, f, zoom) { mode === 'photo-dim' ? 0.34 : 1); } - const r = renderFrame(f, mode === 'oval+photo' ? 'oval' : (showPhoto ? 'off' : mode)); - const img2 = r.toImageData(PALETTE.map((p) => p.hex), zoom); - if (showPhoto) { - // Keep the photo visible wherever the indexed layer is background. - const bg = PALETTE[IDX.bg].hex.replace('#', ''); - const br = parseInt(bg.slice(0, 2), 16), bgn = parseInt(bg.slice(2, 4), 16), bb = parseInt(bg.slice(4, 6), 16); - const d = img2.data; - for (let i = 0; i < d.length; i += 4) { - if (d[i] === br && d[i + 1] === bgn && d[i + 2] === bb) d[i + 3] = 0; - } - const tmp = document.createElement('canvas'); - tmp.width = img2.width; tmp.height = img2.height; - tmp.getContext('2d').putImageData(img2, 0, 0); - g.drawImage(tmp, 0, 0); - } else { - g.putImageData(img2, 0, 0); + blitIndexed(g, renderFrame(f, mode === 'oval+photo' ? 'oval' : (showPhoto ? 'off' : mode)), + zoom, showPhoto); +} + +// Put an indexed raster onto a 2D context. With `keyBg`, background pixels go +// transparent instead of opaque, so whatever was painted underneath - a +// registered photograph, usually - stays visible through them. +// +// Split out of compositeRender so the paint canvas can draw the same pixels at +// a reduced globalAlpha. Nothing else is different about that path, which is +// what keeps the editing view honest: you are dimming the render, not looking +// at a second renderer that might disagree with it. +function blitIndexed(g, raster, zoom, keyBg) { + const img = raster.toImageData(PALETTE.map((p) => p.hex), zoom); + if (!keyBg) { g.putImageData(img, 0, 0); return; } + const bg = PALETTE[IDX.bg].hex.replace('#', ''); + const br = parseInt(bg.slice(0, 2), 16), bgn = parseInt(bg.slice(2, 4), 16), bb = parseInt(bg.slice(4, 6), 16); + const d = img.data; + for (let i = 0; i < d.length; i += 4) { + if (d[i] === br && d[i + 1] === bgn && d[i + 2] === bb) d[i + 3] = 0; } + const tmp = document.createElement('canvas'); + tmp.width = img.width; tmp.height = img.height; + tmp.getContext('2d').putImageData(img, 0, 0); + g.drawImage(tmp, 0, 0); } const keptSorted = () => [...state.keep].sort((a, b) => a - b); @@ -347,11 +626,28 @@ const keptSorted = () => [...state.keep].sort((a, b) => a - b); // Positive lead = the mouth arrives earlier. Only performance parts shift; the // head stays with the audio, because it is the mouth that should anticipate. function leadIndex(f) { - // Reads a cached scalar, not opts(): this runs once per strip thumbnail, and + // Reads cached scalars, not opts(): this runs once per strip thumbnail, and // calling opts() here meant ~14 DOM reads x 74 frames on every redraw. - return shiftIndex(f, state.lead, state.dense.length); + // + // Exposure first, then lead. The grid decides WHICH frames get a new drawing; + // the lead then shifts which pose that drawing carries, by whole frames of the + // original track. Applying them the other way round would put the changes on + // the wrong beats - the picture would update on the odd frames instead of + // holding on the twos. + return shiftIndex(exposeIndex(f, state.exposure), state.lead, state.dense.length); } +// The plate rides the same grid, so the whole picture updates together. On 2s +// means on 2s - a head that cut on the odd frames while the mouth cut on the +// even ones would read as two performances laid over each other. +const plateIndex = (f) => heldFrame(keptSorted(), exposeIndex(f, state.exposure)); + +// Performance tracks that do not take the mouth lead still ride the grid. This +// existing as a named thing is what stopped the eyes holding on 1s in the +// preview while the export held them on 2s - a preview that disagrees with the +// export is the one bug this tool cannot afford. +const perfIndex = (f) => exposeIndex(f, state.exposure); + function blit(canvas, raster, zoom) { canvas.width = RW * zoom; canvas.height = RH * zoom; canvas.getContext('2d').putImageData(raster.toImageData(PALETTE.map((p) => p.hex), zoom), 0, 0); @@ -362,6 +658,7 @@ function drawAll() { drawStrip(); drawWorksheet(); drawReadout(); + drawPaint(); } function drawReadout() { @@ -372,21 +669,43 @@ function drawReadout() { el('readout').textContent = `${state.dense.length} frames → ${kept.length} drawings · ` + `teeth on ${teethFrames}f · ` + + `${blinkRuns(state.eyes.shutR).length}/${blinkRuns(state.eyes.shutL).length} blinks R/L · ` + + `${gazeCells(state.eyes.gaze)} gaze cells · ` + + `${gazeCells(state.brows.R.q)} brow poses · ` + + (state.exposure > 1 + ? `on ${state.exposure}s = ${(state.fps / state.exposure).toFixed(4).replace(/\.?0+$/, '')}fps · ` + : '') + (lead ? `mouth leads ${lead}f (${(lead / state.fps * 1000).toFixed(0)}ms) · ` : '') + `holds ${Math.min(...runs)}–${Math.max(...runs)} frames · ` + `neutral f${state.neutral} · residual ` + `${(state.stab.residual.reduce((a, b) => a + b, 0) / state.dense.length).toFixed(4)}`; } +// Blinks as RUNS, not as shut frames: a three-frame blink is one blink, and the +// count is only useful as "did the performer blink six times or sixty". +function blinkRuns(shut) { + const runs = []; + for (let f = 0; f < shut.length; f++) { + if (shut[f] && !shut[f - 1]) runs.push(f); + } + return runs; +} + +// How many distinct positions the iris ever occupies. This is the number the +// gaze knobs exist to control: two or three is a character who looks at things, +// forty is an unquantised iris sliding around, which is what the grid is for. +const gazeCells = (gaze) => new Set(gaze.map((g) => `${g.x},${g.y}`)).size; + function drawPanes() { const f = state.frame, kept = keptSorted(); - const pf = heldFrame(kept, f); + const pf = plateIndex(f); // The mouth frame is always shown, not only when shifted, so the number can be // watched diverging from f rather than taken on trust. const lead = state.lead; el('framelabel').textContent = `f ${f} / ${state.dense.length - 1} · ${(f / state.fps).toFixed(2)}s · ` + `plate f${pf} · mouth f${leadIndex(f)}` + + (state.exposure > 1 && f % state.exposure ? ' (held)' : '') + (lead ? ` (${lead > 0 ? '+' : ''}${lead} = ${(lead / state.fps * 1000).toFixed(0)}ms)` : '') + (state.keep.has(f) ? ' · KEPT' : ' · held'); @@ -404,12 +723,14 @@ function drawPanes() { const map = (p) => ({ x: dx + (p.x * im.naturalWidth - sx) * s, y: dy + (p.y * im.naturalHeight - sy) * s }); strokePts(g1, LIPS_OUTER.map((i) => map(state.dense[f][i])), '#4ade80'); strokePts(g1, LIPS_INNER.map((i) => map(state.dense[f][i])), '#f87171'); + drawEyeOverlay(g1, map, f); } else { g1.fillStyle = '#555'; g1.font = '13px system-ui'; g1.fillText('synthetic — no source frames', 14, 24); const sc = (p) => ({ x: p.x * c1.width, y: p.y * c1.height }); strokePts(g1, LIPS_OUTER.map((i) => sc(state.dense[f][i])), '#4ade80'); strokePts(g1, LIPS_INNER.map((i) => sc(state.dense[f][i])), '#f87171'); + drawEyeOverlay(g1, sc, f); } const c2 = el('cv-stab'), g2 = c2.getContext('2d'); @@ -426,9 +747,18 @@ function drawPanes() { if (mf !== f) strokePts(g2, z(state.outer[f]), '#2f6b46'); strokePts(g2, z(state.outer[mf]), '#4ade80'); if (!state.hidden[mf]) strokePts(g2, z(state.inner[mf]), '#f87171'); + for (const e of [state.eyes.frames[f].r, state.eyes.frames[f].l]) { + strokePts(g2, z(e.lid), e.shut ? '#f87171' : '#60a5fa'); + if (e.shut) continue; + g2.strokeStyle = '#fbbf24'; + g2.beginPath(); + g2.arc(e.iris.x * ZOOM, e.iris.y * ZOOM, e.iris.r * ZOOM, 0, Math.PI * 2); + g2.stroke(); + } compositeRender(el('cv-render'), f, ZOOM); drawInteriorDebug(f); + drawGazeDebug(f); } // What the teeth measurement actually saw: sampled region, pixels above @@ -458,6 +788,81 @@ function drawInteriorDebug(fRaw) { `area ${m.area}px · ${te.show ? 'SHOWN' : 'hidden'}`; } +// Lid rings and the iris, on the raw frame. Landmark overlays are how you tell +// a tracking failure from a knob set wrong, and the eyes need it more than the +// mouth does: an iris that has latched onto an eyebrow looks, in the flat +// render alone, exactly like a gaze gain that is too high. +function drawEyeOverlay(g, map, f) { + const lm = state.dense[f]; + for (const ring of [EYE_R_RING, EYE_L_RING]) { + strokePts(g, ring.map((i) => map(lm[i])), '#60a5fa'); + } + for (const ring of [BROW_A_RING, BROW_B_RING]) { + strokePts(g, ring.map((i) => map(lm[i])), '#c084fc'); + } + if (!state.eyes.hasIris) return; + for (const iris of [IRIS_A, IRIS_B]) { + strokePts(g, iris.slice(1).map((i) => map(lm[i])), '#fbbf24'); + } +} + +// The gaze field: every position the iris takes over the whole take, plus where +// it is now. Tune against this, not against the numbers - "4 cells" tells you +// the quantisation is working, but only the picture tells you whether the four +// are the four looks the performance actually has. +function drawGazeDebug(f) { + const cv = el('cv-gaze'), S = 150; + cv.width = S; cv.height = S; + const g = cv.getContext('2d'); + g.fillStyle = '#0d0f16'; g.fillRect(0, 0, S, S); + + const ex = state.eyes; + // Scale so the widest excursion in the take fills the box, with a floor so a + // nearly-still gaze does not get magnified into a light show. + let m = 2; + for (const p of ex.gazePx) m = Math.max(m, Math.abs(p.x), Math.abs(p.y)); + const k = (S / 2 - 8) / m; + const X = (v) => S / 2 + v * k, Y = (v) => S / 2 + v * k; + + const o = opts(); + if (o.gazeStep > 0) { + g.strokeStyle = '#1b2030'; g.lineWidth = 1; + for (let i = -20; i <= 20; i++) { + const v = i * o.gazeStep; + if (Math.abs(v) > m) continue; + g.beginPath(); g.moveTo(X(v), 0); g.lineTo(X(v), S); g.stroke(); + g.beginPath(); g.moveTo(0, Y(v)); g.lineTo(S, Y(v)); g.stroke(); + } + } + g.strokeStyle = '#2a2f3e'; + g.beginPath(); g.moveTo(S / 2, 0); g.lineTo(S / 2, S); + g.moveTo(0, S / 2); g.lineTo(S, S / 2); g.stroke(); + + g.fillStyle = '#2f6b46'; + for (const p of ex.gaze) g.fillRect(X(p.x) - 1.5, Y(p.y) - 1.5, 3, 3); + + const raw = ex.gazePx[f], q = ex.gaze[f]; + g.fillStyle = '#8891a5'; + g.fillRect(X(raw.x) - 1, Y(raw.y) - 1, 2, 2); + g.fillStyle = '#fbbf24'; + g.beginPath(); g.arc(X(q.x), Y(q.y), 4, 0, Math.PI * 2); g.fill(); + + const sig = state.eyeSig, fr = state.eyes.frames[f]; + const og = state.gazeOriginValue; + // Per-eye raw gaze is the diagnostic for a wrong-looking eyeline. If the two + // agree and both point the wrong way, the ORIGIN is wrong. If they disagree in + // a sustained way, it is out-of-plane head rotation biasing the projection, + // which no 2D measurement can undo. + const sgn = (v) => `${v >= 0 ? '+' : ''}${v.toFixed(3)}`; + el('eyeinfo').textContent = + `open R ${sig.openR[f].toFixed(3)} L ${sig.openL[f].toFixed(3)} / cut ${o.blinkCut.toFixed(3)}\n` + + `${fr.r.shut ? 'R SHUT ' : ''}${fr.l.shut ? 'L SHUT' : ''}${!fr.r.shut && !fr.l.shut ? 'both open' : ''}\n` + + `gaze ${q.x >= 0 ? '+' : ''}${q.x.toFixed(1)}, ${q.y >= 0 ? '+' : ''}${q.y.toFixed(1)} px\n` + + `raw R ${sgn(sig.gazeR[f].x)} L ${sgn(sig.gazeL[f].x)} (x, eye widths)\n` + + `origin ${o.gazeOrigin} ${sgn(og.x)}, ${sgn(og.y)}` + + (ex.hasIris ? '' : ' — no iris landmarks'); +} + function strokePts(g, pts, color, lw = 1) { g.strokeStyle = color; g.lineWidth = lw; g.beginPath(); @@ -497,7 +902,13 @@ function drawStrip() { const tag = document.createElement('span'); tag.textContent = f; + // Mark the frames that actually carry a drawing. Without it the only way to + // know where your cels are is to scrub and look, and "copy previous" then + // reaches back to somewhere you cannot see. + if ((state.cels.get(f) || []).length) cell.classList.add('cel'); cell.append(cv, tag); + cell.draggable = true; + cell.ondragstart = (ev) => ev.dataTransfer.setData('text/plain', String(f)); cell.onclick = (ev) => { seekTo(f); if (ev.shiftKey) toggle(f); @@ -538,12 +949,67 @@ function drawWorksheet() { /* ---------- export ---------- */ +// Six parts, three per eye, mirroring the mouth's lash/interior/content stack. +// `clip` is what tells the renderer the iris is stencilled by the sclera rather +// than merely drawn after it - without it an extreme gaze would put the iris on +// the cheek. +function eyeParts(grid) { + const out = []; + // Eyes ride the exposure grid but NOT the mouth lead: the lead is a lip-sync + // device and nothing about a blink is tied to the audio. + const src = perfIndex; + [['r', 20], ['l', 23]].forEach(([side, z]) => { + const at = (f) => state.eyes.frames[src(f)][side]; + out.push( + { name: `eye_${side}`, kind: 'poly', z, color: 'skin_dark', interp: 'hold', + keys: grid.map((f) => ({ f, src: src(f), pts: at(f).lash })) }, + { name: `eye_${side}_in`, kind: 'poly', z: z + 1, color: 'eye_white', interp: 'hold', + parent: `eye_${side}`, + keys: grid.map((f) => + (at(f).shut ? { f, hidden: true } : { f, src: src(f), pts: at(f).lid })) }, + { name: `iris_${side}`, kind: 'disc', z: z + 2, color: 'iris', interp: 'hold', + parent: `eye_${side}_in`, clip: `eye_${side}_in`, + keys: grid.map((f) => { + const e = at(f); + return e.shut ? { f, hidden: true } + : { f, src: src(f), c: { x: e.iris.x, y: e.iris.y }, r: e.iris.r }; + }) }, + ); + if (!state.eyes.frames[0].r.pupil) return; + out.push( + { name: `pupil_${side}`, kind: 'rect', z: z + 3, color: 'pupil', interp: 'hold', + parent: `iris_${side}`, clip: `iris_${side}`, + keys: grid.map((f) => { + const e = at(f); + return e.shut ? { f, hidden: true } + : { f, src: src(f), c: { x: e.iris.x, y: e.iris.y }, size: e.pupil }; + }) }, + ); + }); + return out; +} + +// Brows are a traced ring like the lids, so they keep every frame on the grid. +// The quantised raise is already baked into the points - the renderer is handed +// a polygon, not a shape plus an offset it would have to recombine. +function browParts(grid) { + return [['r', 26, 'R'], ['l', 27, 'L']].map(([name, z, side]) => ({ + name: `brow_${name}`, kind: 'poly', z, color: 'brow', interp: 'hold', + keys: grid.map((f) => ({ f, src: perfIndex(f), pts: state.brows[side].frames[perfIndex(f)] })), + })); +} + function exportTake() { const kept = keptSorted(); const N = state.dense.length; + // Output frames that actually carry a key. Everything between them is a hold, + // which the take format already expresses, so on 2s emits half the keys rather + // than emitting each pose twice. + const grid = []; + for (let f = 0; f < N; f += state.exposure) grid.push(f); const take = { name: el('takename').value || 'line_01', - frames: N, width: RW, height: RH, exposure: 1, fps: state.fps, + frames: N, width: RW, height: RH, exposure: state.exposure, fps: state.fps, palette: PALETTE, slot: { x: RW / 2, y: RH / 2 }, parts: [ @@ -554,14 +1020,20 @@ function exportTake() { // key f carries the pose from source frame f+lead - so the renderer never // needs to know about it. { name: 'mouth', kind: 'poly', z: 30, color: 'skin_dark', interp: 'hold', - keys: state.outer.map((_, f) => ({ f, src: leadIndex(f), pts: state.outer[leadIndex(f)] })) }, + keys: grid.map((f) => ({ f, src: leadIndex(f), pts: state.outer[leadIndex(f)] })) }, { name: 'mouth_in', kind: 'poly', z: 31, color: 'mouth_dark', interp: 'hold', parent: 'mouth', - keys: state.inner.map((_, f) => { + keys: grid.map((f) => { const m = leadIndex(f); return state.hidden[m] ? { f, hidden: true } : { f, src: m, pts: state.inner[m] }; }) }, + // Eyes. The lids are traced, so like the mouth they cost nothing and keep + // every frame. The iris is a primitive: its quantised position means the + // key stream is dense but the VALUES change only on saccades, so a + // hold-interpolating renderer cuts between fixations by itself. + ...eyeParts(grid), + ...browParts(grid), { name: 'teeth', kind: 'poly', z: 32, color: 'teeth', interp: 'hold', parent: 'mouth_in', - keys: state.teeth.map((_, f) => { + keys: grid.map((f) => { const m = leadIndex(f), te = state.teeth[m]; return te.show && te.pts ? { f, src: m, pts: te.pts } : { f, hidden: true }; }) }, @@ -577,7 +1049,9 @@ function exportTake() { a.href = URL.createObjectURL(new Blob([text], { type: 'text/plain' })); a.download = `${take.name}.take`; a.click(); - status(`exported — ${kept.length} plate drawings, ${N} mouth frames`, 'ok'); + status(`exported — ${kept.length} plate drawings, ${grid.length} mouth keys` + + (state.exposure > 1 ? ` on ${state.exposure}s` : '') + ', ' + + `${blinkRuns(state.eyes.shutR).length + blinkRuns(state.eyes.shutL).length} blinks`, 'ok'); } /* ---------- wiring ---------- */ @@ -605,6 +1079,8 @@ async function runFrames() { state.interior = measureAll(images, dense, opts()); el('scrub').max = dense.length - 1; state.frame = 0; + labelExposure(); + loadCels(); rebuild(true); const dur = (dense.length / state.fps).toFixed(2); status(`${images.length} frames · ${images[0].naturalWidth}x${images[0].naturalHeight} · ` + @@ -626,24 +1102,49 @@ function runSynthetic() { state.dense = synthDense(72); el('scrub').max = 71; state.frame = 0; + labelExposure(); + loadCels(); rebuild(true); status('synthetic — exercises everything below detection', 'ok'); } +// How each slider's raw value reads out. A table rather than the conditional +// chain this used to be: that chain grew a branch per knob and was one ternary +// away from being unreadable. +const FMT = { + apertureThresh: (v) => (v / 1000).toFixed(3), + tol: (v) => (v / 1000).toFixed(3), + blinkCut: (v) => (v / 1000).toFixed(3), + teethOn: (v) => (v / 100).toFixed(2), + teethErode: (v) => (v / 100).toFixed(2), + tongueReject: (v) => (v / 100).toFixed(2), + topBias: (v) => (v / 100).toFixed(2), + irisSize: (v) => `${v}%`, + gazeGain: (v) => (v / 100).toFixed(2), + gazeStep: (v) => (v ? `${v}px` : 'off'), + browStep: (v) => (v ? `${v}px` : 'off'), + browWeight: (v) => `${v}px`, + browGain: (v) => (v / 100).toFixed(2), + pupilPx: (v) => (v ? `${v}px` : 'off'), + lashPx: (v) => `${v}px`, + lead: (v) => (v > 0 ? `+${v}` : String(v)), +}; + for (const id of ['verts', 'smoothWin', 'contourSmooth', 'apertureThresh', 'tol', 'teethOn', 'teethDwell', 'teethErode', 'tongueReject', 'blobGrow', - 'topBias', 'teethVerts', 'teethSmooth', 'lead']) { + 'topBias', 'teethVerts', 'teethSmooth', 'lead', + 'eyeVerts', 'lashPx', 'irisSize', 'pupilPx', 'gazeGain', + 'gazeStep', 'gazeDwell', 'blinkCut', 'blinkHold', 'blinkDwell', + 'browVerts', 'browWeight', 'browGain', 'browStep', 'browDwell']) { + const show = () => { + el(id + 'v').textContent = FMT[id] ? FMT[id](+el(id).value) : el(id).value; + }; el(id).addEventListener('input', () => { - el(id + 'v').textContent = id === 'apertureThresh' || id === 'tol' - ? (+el(id).value / 1000).toFixed(3) - : ['teethOn', 'teethErode', 'tongueReject', 'topBias'].includes(id) - ? (+el(id).value / 100).toFixed(2) - : id === 'lead' && +el(id).value > 0 ? `+${el(id).value}` - : el(id).value; + show(); if (id === 'tol') return; // tol only matters when you ask for a suggestion rebuild(false); }); - el(id + 'v').textContent = el(id).value; + show(); } function seekTo(f) { @@ -657,6 +1158,22 @@ el('btn-frames').onclick = runFrames; el('btn-synth').onclick = runSynthetic; el('btn-export').onclick = exportTake; el('plateMode').addEventListener('change', () => { if (state.dense) drawAll(); }); +el('exposure').addEventListener('change', () => { if (state.dense) rebuild(false); }); +for (const id of ['irisAnchor', 'gazeOrigin']) { + el(id).addEventListener('change', () => { if (state.dense) rebuild(false); }); +} + +// Label the exposure options in the only units that mean anything here: the +// rate the picture actually changes at, which depends on the clip's own rate. +// "on 2s" is the animator's name for it and the number is what you hear against +// the audio, so the menu says both. +function labelExposure() { + for (const opt of el('exposure').options) { + const n = +opt.value; + const rate = (state.fps / n).toFixed(4).replace(/\.?0+$/, ''); + opt.textContent = `${rate} fps · on ${n}s`; + } +} el('btn-saveframe').onclick = () => { if (!state.dense) return; const cv = document.createElement('canvas'); @@ -763,6 +1280,144 @@ PALETTE.forEach((p) => { el('palette').append(sw); }); +/* ---------- paint ---------- */ + +// The cel being edited is the one ON SCREEN, which is the most recent kept +// frame at or before the playhead. You can scrub anywhere and keep drawing on +// the cel you can see, rather than having to land exactly on a kept frame. +const celFrame = () => (state.dense ? plateIndex(state.frame) : 0); + +const paint = new PaintUI({ + canvas: el('cv-paint'), list: el('paintlayers'), info: el('paintinfo'), + zoom: 3, RW, RH, palette: PALETTE, + getCel: () => (state.dense ? (state.cels.get(celFrame()) || []) : []), + setCel: (cel) => { if (state.dense) state.cels.set(celFrame(), cel); }, + celAt: (f) => (state.dense ? state.cels.get(heldFrame(keptSorted(), f)) : null), + backing: (cv, z) => { + if (!state.dense) { + cv.width = RW * z; cv.height = RH * z; + const g = cv.getContext('2d'); + g.fillStyle = '#0d0f16'; g.fillRect(0, 0, cv.width, cv.height); + return; + } + const op = +el('paintOpacity').value / 100; + // At 100% this is exactly the normal composite, so the slider changes + // nothing at all until you reach for it. + if (op >= 1) { compositeRender(cv, state.frame, z); return; } + paintGhost(cv, z, op); + }, + // paintLabels, not drawPaint: drawPaint re-renders the canvas, and this runs + // from inside commit(), which renders immediately afterwards anyway. + onChange: () => { saveCels(); paintLabels(); drawPanes(); drawStrip(); drawWorksheet(); }, +}); + +el('paintTool').addEventListener('change', () => { + paint.tool = el('paintTool').value; + paint.draft = null; + paint.render(); +}); +el('paintColor').addEventListener('change', () => { paint.color = +el('paintColor').value; }); +el('paintOpacity').addEventListener('input', () => { + el('paintOpacityv').textContent = `${el('paintOpacity').value}%`; + paint.render(); +}); +el('paintOpacityv').textContent = `${el('paintOpacity').value}%`; +el('btn-celclear').onclick = () => { + if (!state.dense) return; + state.cels.set(celFrame(), newCel()); + paint.sel = -1; + paint.commit(); +}; +el('btn-celprev').onclick = () => { + // Copy the last finished drawing onto this one - the case you reach for + // constantly, stepping forward and carrying the previous cel with you. + // + // "The last drawing" means the nearest earlier enabled frame that ACTUALLY + // HAS one, not simply the nearest earlier enabled frame. Every frame is + // enabled until you curate the strip, so the naive rule resolved to f-1, + // which is empty, and the button looked like it only ever copied the frame + // immediately to the left. Skipping the empties makes it behave the same + // before and after you thin the strip out. + if (!state.dense) return; + const here = celFrame(); + const prev = keptSorted() + .filter((k) => k < here && (state.cels.get(k) || []).length) + .pop(); + if (prev === undefined) { paint.say('no earlier drawing to copy'); return; } + state.cels.set(here, cloneCel(state.cels.get(prev))); + paint.sel = -1; + paint.commit(); + paint.say(`copied f${prev} → f${here} — ${state.cels.get(here).length} layers`); +}; + +PALETTE.forEach((p, i) => { + const o = document.createElement('option'); + o.value = i; o.textContent = p.name; + el('paintColor').append(o); +}); +el('paintColor').value = 1; +paint.color = 1; + +// Drawings are the only thing here a person made by hand, so losing them to a +// reload would be the worst failure in the tool. Everything else regenerates. +const celKey = () => `arthur.cels.${el('takename').value || 'line_01'}`; +function saveCels() { + try { + localStorage.setItem(celKey(), JSON.stringify([...state.cels])); + } catch { /* private window, quota - not worth failing a brush stroke over */ } +} +function loadCels() { + state.cels = new Map(); + try { + const raw = localStorage.getItem(celKey()); + if (raw) state.cels = new Map(JSON.parse(raw).map(([k, v]) => [+k, v])); + } catch { /* corrupt or absent: start empty */ } +} + +function drawPaint() { + paintLabels(); + paint.render(); +} + +// The drawing dimmed over the source frame, for tracing. +// +// An EDITING AID ONLY - opacity is never written to the raster, never exported, +// and the flat render is untouched. Nothing here may reach the output: a +// translucent fill is the one thing this format cannot express, so if it ever +// leaked into the rasteriser it would have to be flattened against a background +// and would silently become a colour that is not in the ramp. +// +// The plate oval is deliberately not drawn. It is a stand-in for art that does +// not exist yet, and it covers the performer's face in flat skin - which is +// exactly the part of the frame you turned the opacity down to look at. +function paintGhost(cv, z, op) { + cv.width = RW * z; cv.height = RH * z; + const g = cv.getContext('2d'); + g.fillStyle = PALETTE[IDX.bg].hex; + g.fillRect(0, 0, cv.width, cv.height); + + const f = state.frame, pf = plateIndex(f); + const img = state.images[pf]; + if (img) drawRegistered(g, img, state.stab.transforms[pf], state.xform, z, 1); + + g.globalAlpha = op; + blitIndexed(g, renderFrame(f, 'off'), z, !!img); + g.globalAlpha = 1; +} + +function paintLabels() { + if (!state.dense) return; + const kept = keptSorted(), cf = celFrame(); + const until = (kept[kept.indexOf(cf) + 1] ?? state.dense.length) - 1; + const drawn = kept.filter((k) => (state.cels.get(k) || []).length); + el('paintframe').textContent = + `drawing cel f${cf}` + (until > cf ? ` — holds to f${until}` : '') + + (state.frame !== cf ? ` · playhead f${state.frame}` : ''); + el('paintcels').textContent = drawn.length + ? `${drawn.length} drawn: ${drawn.join(' ')}` + : 'nothing drawn yet'; +} + // #synth / #frames autorun, so the tool can be driven headlessly for smoke tests // and deep-linked. Detection needs WebGL; the synthetic path does not. window.addEventListener('error', (e) => { @@ -775,3 +1430,6 @@ else if (location.hash === '#frames') runFrames(); else status('ready — Load frames, then step with \u2190 \u2192 and delete with X'); window.__roto = state; // headless smoke test reads this +window.__render = compositeRender; // ...and renders arbitrary frames off-screen +window.__lead = leadIndex; // ...and resolves the performance frame +window.__drawAll = drawAll; // ...and forces a full redraw diff --git a/js/landmarks.js b/js/landmarks.js index 8251e6b..e8e7b4a 100644 --- a/js/landmarks.js +++ b/js/landmarks.js @@ -52,3 +52,77 @@ export function subsampleSlots(len, n) { export function subsampleRing(ring, n) { return subsampleSlots(ring.length, n).map((s) => ring[s]); } + +// ---- eyes ---- +// +// Eyelid rings, under the same contract as the lip rings: ORDERED traversals +// where slot position IS vertex identity. Both eyes start at the OUTER corner +// and go over the UPPER lid first, so slot k means the same anatomy on both +// sides. On a 16-slot ring that puts the four cardinals exactly on the four +// quarter slots - 0 outer corner, 4 upper lid centre, 8 inner corner, 12 lower +// lid centre - so every even vertex budget lands on real landmarks. +// +// The two rings traverse opposite directions on screen, because they are +// mirrored anatomy described the same way. Nothing downstream cares: an +// even-odd fill has no winding, and ring SIMPLICITY is what is asserted. +export const EYE_R_RING = [ + 33, 246, 161, 160, 159, 158, 157, 173, + 133, 155, 154, 153, 145, 144, 163, 7, +]; +export const EYE_L_RING = [ + 263, 466, 388, 387, 386, 385, 384, 398, + 362, 382, 381, 380, 374, 373, 390, 249, +]; + +// Outer, inner corner per eye. All four are also in RIGID, and that is the +// point: the eye's reference frame is built only from landmarks that do not +// move under performance, so a blink cannot be mistaken for a change of gaze. +export const EYE_R_CORNERS = [33, 133]; +export const EYE_L_CORNERS = [263, 362]; + +// Upper and lower lid centres. Their separation over the corner distance is the +// openness signal that decides whether the eye is shut - the same shape of +// measurement as APERTURE is for the mouth, but normalised, so one threshold +// carries across takes and faces. +export const EYE_R_LIDS = [159, 145]; +export const EYE_L_LIDS = [386, 374]; + +// The two iris blocks the refined mesh appends: centre first, then four ring +// points. WHICH BLOCK BELONGS TO WHICH EYE IS NOT DECLARED HERE - MediaPipe's +// own "left"/"right" is viewer-relative in some docs and subject-relative in +// others, and a swap looks almost right, so it would survive an eyeball and +// then read as a permanently wall-eyed character. pipeline.js resolves it from +// the geometry instead. +export const IRIS_A = [468, 469, 470, 471, 472]; +export const IRIS_B = [473, 474, 475, 476, 477]; + +// ---- brows ---- +// +// Each brow is two five-point chains, an upper edge and a lower edge, which +// close into a ten-point ring: out along one edge from the outer end to the +// inner, back along the other. +// +// WHICH EDGE IS UPPER IS DELIBERATELY NOT DECLARED, and unlike the iris it does +// not need to be. Swapping them traverses the same ring the other way round, +// and an even-odd fill has no winding, so the shape is identical either way. +// What the ring guarantees instead is that the two ENDS land on fixed slots: +// 0 and 9 are one end, 4 and 5 the other. Averaging a pair therefore gives the +// brow's height at that end whichever edge is on top, which is all the raise +// and tilt measurement needs. +// +// Which end is the OUTER one is resolved from geometry in pipeline.js, because +// getting it backwards mirrors the tilt - inner-up "worried" would render as +// outer-up - and that is a expression error, not a glitch, so it would read as +// a directed performance choice rather than as a bug. +export const BROW_A_RING = [ + 70, 63, 105, 66, 107, + 55, 65, 52, 53, 46, +]; +export const BROW_B_RING = [ + 300, 293, 334, 296, 336, + 285, 295, 282, 283, 276, +]; + +// The slots at each end of a brow ring, as pairs to average. +export const BROW_END_0 = [0, 9]; +export const BROW_END_1 = [4, 5]; diff --git a/js/mathutil.js b/js/mathutil.js index 310eb8c..b49c51c 100644 --- a/js/mathutil.js +++ b/js/mathutil.js @@ -108,3 +108,27 @@ export function smoothTransforms(tfs, radius) { theta: Math.atan2(sn[i], c[i]), s: s[i], tx: tx[i], ty: ty[i], })); } + +// Push a ring outward from its centroid by a FIXED distance, not by a scale +// factor. +// +// Scaling collapses with the shape: a shut eyelid scaled by 1.1 is still a shut +// eyelid, so the lash line - the only thing left to draw when the eye is closed +// - would vanish exactly on the frames where it is the whole drawing. A fixed +// radial offset gives a band of roughly constant thickness that survives the +// ring going degenerate, and it keeps a star-shaped ring simple, which +// docs/design.md requires of every cut part. +export function offsetRing(pts, d) { + if (!d) return pts; + let cx = 0, cy = 0; + for (const p of pts) { cx += p.x; cy += p.y; } + cx /= pts.length; cy /= pts.length; + return pts.map((p) => { + const dx = p.x - cx, dy = p.y - cy; + const m = Math.hypot(dx, dy); + // A vertex sitting exactly on the centroid has no outward direction. Leave + // it where it is rather than emitting NaN and poisoning the whole ring. + return m < 1e-9 ? { x: p.x, y: p.y } + : { x: p.x + (dx / m) * d, y: p.y + (dy / m) * d }; + }); +} diff --git a/js/paint.js b/js/paint.js new file mode 100644 index 0000000..b50648a --- /dev/null +++ b/js/paint.js @@ -0,0 +1,345 @@ +// Vector cel painting: flat polygons on the frames that get their own drawing. +// +// THIS IS A SKETCH. It exists to test whether the aesthetic holds when a human +// draws the background rather than the tracker deriving it, and it is expected +// to be replaced by a real paint surface with onion skin, undo and a proper +// tool model. Deliberately kept to one module with no dependencies on the rest +// of the pipeline so that throwing it away is a delete rather than a surgery. +// +// The data model is the part worth keeping: +// +// cel = Layer[] index 0 is the BACK of the stack, last is the front +// Layer = { name, color, pts, hidden } +// +// `color` is a PALETTE INDEX, never an RGB value - the one rule from +// docs/design.md that a paint tool could most easily break. You cannot pick a +// colour here that is not already in the take's ramp. +// +// Points are integers in 320x200 raster space. Snapping is not a convenience: +// sub-pixel vertices on a hard-edged indexed rasteriser move an edge by a whole +// pixel or not at all depending on where the polygon happens to land, so a +// shape nudged by 0.4px shimmers instead of holding still. + +export const newCel = () => []; + +export const cloneCel = (cel) => + (cel || []).map((l) => ({ + name: l.name, color: l.color, hidden: !!l.hidden, + pts: l.pts.map((p) => ({ x: p.x, y: p.y })), + })); + +// Fill every visible layer back to front. Same rasteriser the derived parts +// use, so a painted shape and a traced one cannot look different. +export function drawCel(raster, cel, idxOf) { + if (!cel) return; + for (const l of cel) { + if (l.hidden || l.pts.length < 3) continue; + raster.fillPoly(l.pts, idxOf(l.color)); + } +} + +const pointInPoly = (pts, x, y) => { + let inside = false; + for (let i = 0, j = pts.length - 1; i < pts.length; j = i++) { + if ((pts[i].y > y) !== (pts[j].y > y) && + x < ((pts[j].x - pts[i].x) * (y - pts[i].y)) / (pts[j].y - pts[i].y) + pts[i].x) inside = !inside; + } + return inside; +}; + +// Distance from p to segment ab, and where along it the foot falls. +function segDist(p, a, b) { + const vx = b.x - a.x, vy = b.y - a.y; + const len = vx * vx + vy * vy; + const t = len ? Math.max(0, Math.min(1, ((p.x - a.x) * vx + (p.y - a.y) * vy) / len)) : 0; + const cx = a.x + t * vx, cy = a.y + t * vy; + return { d: Math.hypot(p.x - cx, p.y - cy), t }; +} + +export class PaintUI { + constructor(opts) { + this.canvas = opts.canvas; // where the cel is drawn and edited + this.list = opts.list; // layer stack DOM host + this.info = opts.info; // status line + this.zoom = opts.zoom || 3; + this.RW = opts.RW; this.RH = opts.RH; + this.palette = opts.palette; // [{name,hex}], index is the colour + this.getCel = opts.getCel; // () => Layer[] for the frame being edited + this.setCel = opts.setCel; // (Layer[]) => void + this.celAt = opts.celAt; // (frame) => Layer[] for drag-to-seed + this.backing = opts.backing; // (canvas, zoom) => void, paints what is under + this.onChange = opts.onChange; // commit + redraw everything else + + this.tool = 'pen'; + this.color = 1; + this.sel = -1; // selected layer index + this.draft = null; // in-progress pen path + this.drag = null; + + this.bind(); + } + + /* ---- geometry ---- */ + + at(ev) { + const r = this.canvas.getBoundingClientRect(); + // Snap to the raster grid. See the note at the top of the file. + return { + x: Math.round(((ev.clientX - r.left) / r.width) * this.RW), + y: Math.round(((ev.clientY - r.top) / r.height) * this.RH), + }; + } + + hit(p) { + const cel = this.getCel(); + const near = 3; + // Vertices of the SELECTED layer win over everything, so a vertex sitting + // under another shape stays grabbable instead of selecting the shape on top. + if (this.sel >= 0 && cel[this.sel]) { + const pts = cel[this.sel].pts; + for (let i = 0; i < pts.length; i++) { + if (Math.hypot(pts[i].x - p.x, pts[i].y - p.y) <= near) return { layer: this.sel, vertex: i }; + } + for (let i = 0; i < pts.length; i++) { + const s = segDist(p, pts[i], pts[(i + 1) % pts.length]); + if (s.d <= near / 2) return { layer: this.sel, edge: i }; + } + } + // Then shapes, front to back, so clicking picks what you can see. + for (let i = cel.length - 1; i >= 0; i--) { + if (!cel[i].hidden && cel[i].pts.length >= 3 && pointInPoly(cel[i].pts, p.x, p.y)) { + return { layer: i }; + } + } + return null; + } + + /* ---- editing ---- */ + + commit() { this.onChange(); this.render(); } + + addLayer(pts) { + const cel = this.getCel().slice(); + cel.push({ name: `shape ${cel.length + 1}`, color: this.color, hidden: false, pts }); + this.setCel(cel); + this.sel = cel.length - 1; + this.commit(); + } + + mutate(fn) { + const cel = cloneCel(this.getCel()); + fn(cel); + this.setCel(cel); + this.commit(); + } + + move(from, to) { + if (to < 0 || to >= this.getCel().length) return; + this.mutate((cel) => { cel.splice(to, 0, cel.splice(from, 1)[0]); }); + this.sel = to; + this.render(); + } + + bind() { + const c = this.canvas; + + c.addEventListener('mousedown', (ev) => { + ev.preventDefault(); + const p = this.at(ev); + + if (this.tool === 'pen') { + if (!this.draft) { this.draft = [p]; this.render(); return; } + const first = this.draft[0]; + // Closing on the first point is how you finish; three points minimum, + // because a two-point "polygon" fills nothing and looks like a bug. + if (this.draft.length >= 3 && Math.hypot(first.x - p.x, first.y - p.y) <= 3) { + const pts = this.draft; this.draft = null; this.addLayer(pts); return; + } + this.draft.push(p); this.render(); return; + } + + const h = this.hit(p); + if (!h) { this.sel = -1; this.render(); return; } + this.sel = h.layer; + + if (h.vertex !== undefined) { + if (ev.altKey) { + // Never below a triangle. + if (this.getCel()[h.layer].pts.length > 3) { + this.mutate((cel) => cel[h.layer].pts.splice(h.vertex, 1)); + } + return; + } + this.drag = { kind: 'vertex', layer: h.layer, vertex: h.vertex }; + this.render(); return; + } + if (h.edge !== undefined && ev.shiftKey) { + this.mutate((cel) => cel[h.layer].pts.splice(h.edge + 1, 0, { x: p.x, y: p.y })); + this.drag = { kind: 'vertex', layer: h.layer, vertex: h.edge + 1 }; + return; + } + this.drag = { kind: 'shape', layer: h.layer, from: p }; + this.render(); + }); + + window.addEventListener('mousemove', (ev) => { + if (!this.drag) return; + const p = this.at(ev); + const d = this.drag; + const cel = this.getCel(); + if (d.kind === 'vertex') { + const v = cel[d.layer].pts[d.vertex]; + if (v.x === p.x && v.y === p.y) return; + v.x = p.x; v.y = p.y; + } else { + const dx = p.x - d.from.x, dy = p.y - d.from.y; + if (!dx && !dy) return; + for (const q of cel[d.layer].pts) { q.x += dx; q.y += dy; } + d.from = p; + } + this.render(); + }); + + window.addEventListener('mouseup', () => { + if (!this.drag) return; + this.drag = null; + this.commit(); + }); + + c.addEventListener('dblclick', (ev) => { + ev.preventDefault(); + if (this.tool === 'pen' && this.draft && this.draft.length >= 3) { + const pts = this.draft; this.draft = null; this.addLayer(pts); + } + }); + + // Seed from another frame: drag a strip thumbnail onto the canvas and take + // a copy of whatever that frame is showing, every layer. Copy, never + // reference - two cels sharing a layer object would edit each other and the + // reason would be invisible. + c.addEventListener('dragover', (ev) => { ev.preventDefault(); c.classList.add('drop'); }); + c.addEventListener('dragleave', () => c.classList.remove('drop')); + c.addEventListener('drop', (ev) => { + ev.preventDefault(); + c.classList.remove('drop'); + const f = +ev.dataTransfer.getData('text/plain'); + if (!Number.isFinite(f)) return; + const src = this.celAt(f); + if (!src || !src.length) { this.say(`f${f} has nothing to copy`); return; } + this.setCel(cloneCel(src)); + this.sel = -1; + this.commit(); + this.say(`seeded from f${f} — ${src.length} layers`); + }); + + window.addEventListener('keydown', (ev) => { + if (ev.target.tagName === 'INPUT' || ev.target.tagName === 'SELECT') return; + if (!this.focused) return; + if (ev.key === 'Escape' && this.draft) { this.draft = null; this.render(); ev.preventDefault(); } + else if (ev.key === 'Enter' && this.draft && this.draft.length >= 3) { + const pts = this.draft; this.draft = null; this.addLayer(pts); ev.preventDefault(); + } else if (ev.key === 'Backspace' && this.draft) { + this.draft.pop(); if (!this.draft.length) this.draft = null; + this.render(); ev.preventDefault(); + } else if ((ev.key === 'Delete' || ev.key === 'Backspace') && this.sel >= 0) { + this.mutate((cel) => cel.splice(this.sel, 1)); + this.sel = -1; ev.preventDefault(); + } + }); + + // The paint canvas takes the keyboard only while the pointer is over it, so + // the main window's frame stepping keeps working everywhere else. + c.addEventListener('mouseenter', () => { this.focused = true; }); + c.addEventListener('mouseleave', () => { this.focused = false; }); + } + + say(msg) { if (this.info) this.info.textContent = msg; } + + /* ---- drawing ---- */ + + render() { + const z = this.zoom, c = this.canvas; + this.backing(c, z); // the composited frame, cels included + const g = c.getContext('2d'); + + const cel = this.getCel(); + g.lineWidth = 1; + cel.forEach((l, i) => { + if (l.pts.length < 2) return; + const on = i === this.sel; + // Unselected outlines stay faint: they are there so you can find a shape + // to click, not so you can read them. + g.strokeStyle = l.hidden ? '#f8717166' : (on ? '#fbbf24' : '#ffffff33'); + g.beginPath(); + l.pts.forEach((p, k) => (k ? g.lineTo(p.x * z, p.y * z) : g.moveTo(p.x * z, p.y * z))); + g.closePath(); g.stroke(); + if (!on) return; + g.fillStyle = '#fbbf24'; + for (const p of l.pts) g.fillRect(p.x * z - 2, p.y * z - 2, 5, 5); + }); + + if (this.draft) { + g.strokeStyle = '#4ade80'; + g.beginPath(); + this.draft.forEach((p, k) => (k ? g.lineTo(p.x * z, p.y * z) : g.moveTo(p.x * z, p.y * z))); + g.stroke(); + g.fillStyle = '#4ade80'; + for (const p of this.draft) g.fillRect(p.x * z - 2, p.y * z - 2, 5, 5); + // The closing target, so "click here to finish" is visible rather than + // something you have to know. + if (this.draft.length >= 3) { + g.strokeStyle = '#4ade80'; + g.strokeRect(this.draft[0].x * z - 4, this.draft[0].y * z - 4, 9, 9); + } + } + this.renderList(); + } + + renderList() { + const cel = this.getCel(); + this.list.innerHTML = ''; + // Top of the list is the FRONT of the stack, the way a layers panel reads. + for (let i = cel.length - 1; i >= 0; i--) { + const l = cel[i]; + const row = document.createElement('div'); + row.className = 'lay' + (i === this.sel ? ' sel' : ''); + + const sw = document.createElement('select'); + this.palette.forEach((p, k) => { + const o = document.createElement('option'); + o.value = k; o.textContent = p.name; + sw.append(o); + }); + sw.value = l.color; + sw.onchange = () => this.mutate((cc) => { cc[i].color = +sw.value; }); + + const chip = document.createElement('i'); + chip.style.background = this.palette[l.color] ? this.palette[l.color].hex : '#f0f'; + + const nm = document.createElement('b'); + nm.textContent = l.name; + nm.onclick = () => { this.sel = i; this.render(); }; + + const btn = (txt, title, fn) => { + const b = document.createElement('button'); + b.textContent = txt; b.title = title; + b.onclick = (e) => { e.stopPropagation(); fn(); }; + return b; + }; + + row.append(chip, sw, nm, + btn('↑', 'forward', () => this.move(i, i + 1)), + btn('↓', 'back', () => this.move(i, i - 1)), + btn(l.hidden ? '◻' : '◼', 'show/hide', () => this.mutate((cc) => { cc[i].hidden = !cc[i].hidden; })), + btn('✕', 'delete', () => { this.mutate((cc) => cc.splice(i, 1)); this.sel = -1; })); + row.onclick = () => { this.sel = i; this.render(); }; + this.list.append(row); + } + if (!cel.length) { + const e = document.createElement('div'); + e.className = 'legend'; + e.textContent = 'no layers — draw with the pen, or drag a frame here to copy its drawing'; + this.list.append(e); + } + } +} diff --git a/js/pipeline.js b/js/pipeline.js index 2e1155c..85bfd30 100644 --- a/js/pipeline.js +++ b/js/pipeline.js @@ -2,7 +2,10 @@ // All policy lives here, never in the renderer. See docs/design.md, // "The take is the contract". -import { RIGID, LIPS_OUTER, LIPS_INNER, APERTURE, FACE_OVAL, EYE_INNER, subsampleSlots } from './landmarks.js'; +import { RIGID, LIPS_OUTER, LIPS_INNER, APERTURE, FACE_OVAL, EYE_INNER, + EYE_R_RING, EYE_L_RING, EYE_R_CORNERS, EYE_L_CORNERS, + EYE_R_LIDS, EYE_L_LIDS, IRIS_A, IRIS_B, + BROW_A_RING, BROW_B_RING, BROW_END_0, BROW_END_1, subsampleSlots } from './landmarks.js'; import { fitSimilarity, applySimAll, applySim, fitResidual, procrustesMean, smoothTransforms, movingAverage } from './mathutil.js'; // MediaPipe normalises x by image WIDTH and y by image HEIGHT, so its normalised @@ -25,6 +28,13 @@ export function stabilize(dense, smoothRadius, aspect = 1) { const raw = rigid.map((r) => fitSimilarity(r, ref)); const tfs = smoothTransforms(raw, smoothRadius); + // The refined mesh appends ten iris points to the 468 face points, but a + // plain mesh does not, and synthetic or hand-fed tracks need not. Checked + // rather than assumed: reading past the end would surface as NaN gaze deep + // downstream instead of as "this track carries no iris". + const hasIris = dense.every((f) => f && f.length > IRIS_B[IRIS_B.length - 1]); + const map = (table) => dense.map((f, i) => applySimAll(tfs[i], pick(f, table, aspect))); + return { ref, transforms: tfs, @@ -43,9 +53,296 @@ export function stabilize(dense, smoothRadius, aspect = 1) { const a = applySimAll(tfs[i], pick(f, APERTURE, aspect)); return Math.hypot(a[0].x - a[1].x, a[0].y - a[1].y); }), + // Eyes. Lid rings are a feature and get traced like the mouth; corners and + // lid centres are the measurement frame; the iris blocks are raw until + // pairIrises decides which is which. + lidR: map(EYE_R_RING), lidL: map(EYE_L_RING), + cornersR: map(EYE_R_CORNERS), cornersL: map(EYE_L_CORNERS), + lidsR: map(EYE_R_LIDS), lidsL: map(EYE_L_LIDS), + irisA: hasIris ? map(IRIS_A) : null, + irisB: hasIris ? map(IRIS_B) : null, + browA: map(BROW_A_RING), browB: map(BROW_B_RING), }; } +/* ---------- brows ---------- */ + +// Two correspondences resolved from geometry, for the same reason the iris +// pairing is: a wrong guess here is survivable enough to escape notice. +// +// Which ring is which brow follows MediaPipe's left/right naming, which is the +// naming that would have put the irises on the wrong eyes. Which END of a ring +// is the OUTER one matters more: get it backwards and the tilt mirrors, so +// inner-up "worried" renders as outer-up, which is a different expression +// rather than a broken one. It would read as a directed performance choice and +// never be questioned. +// +// Both are decided by voting across every frame against landmarks already known +// to be rigid, so one bad detection cannot swing them. +export function pairBrows(stab) { + const N = stab.browA.length; + const cen = (ring) => { + let x = 0; + for (const p of ring) x += p.x; + return x / ring.length; + }; + let side = 0, ends = 0; + for (let f = 0; f < N; f++) { + const cR = mid(stab.cornersR[f][0], stab.cornersR[f][1]).x; + const cL = mid(stab.cornersL[f][0], stab.cornersL[f][1]).x; + side += Math.abs(cen(stab.browA[f]) - cR) < Math.abs(cen(stab.browA[f]) - cL) ? 1 : -1; + + // EYE_R_CORNERS is [outer, inner], so this asks whether slot 0 of the ring + // sits nearer the eye's outer corner than its inner one. + const ring = side > 0 ? stab.browA[f] : stab.browB[f]; + const co = side > 0 ? stab.cornersR[f] : stab.cornersL[f]; + const s0 = ring[BROW_END_0[0]]; + ends += Math.abs(s0.x - co[0].x) < Math.abs(s0.x - co[1].x) ? 1 : -1; + } + return { + right: side > 0 ? 'browA' : 'browB', + left: side > 0 ? 'browB' : 'browA', + outerAtSlot0: ends > 0, + }; +} + +// Brow height above its own eye, at each end, in eye widths. +// +// Measured against the eye's CORNER MIDPOINT, not the lid: the corners are +// rigid, so a blink cannot read as a brow raise. That is the same trap the gaze +// origin has and it is worth avoiding twice - brows and lids move together +// constantly, and a brow that jumped on every blink would look like a tic. +// +// Two ends rather than one height, because raise and tilt are different +// expressions built from the same measurement: both ends up is surprise, inner +// up alone is worry, inner down is anger. One number could not tell them apart. +export function browSignals(stab) { + const N = stab.browA.length; + const pairing = pairBrows(stab); + const endOuter = pairing.outerAtSlot0 ? BROW_END_0 : BROW_END_1; + const endInner = pairing.outerAtSlot0 ? BROW_END_1 : BROW_END_0; + const out = { R: [], L: [], pairing }; + + for (let f = 0; f < N; f++) { + for (const [side, corners] of [['R', stab.cornersR], ['L', stab.cornersL]]) { + const ring = stab[pairing[side === 'R' ? 'right' : 'left']][f]; + const c = mid(corners[f][0], corners[f][1]); + const w = dist(corners[f][0], corners[f][1]); + const at = (pair) => (ring[pair[0]].y + ring[pair[1]].y) / 2; + // y grows downward, so a brow ABOVE the eye gives a positive raise. + out[side].push({ x: (c.y - at(endOuter)) / w, y: (c.y - at(endInner)) / w }); + } + } + return out; +} + +/* ---------- eyes ---------- */ + +const mid = (a, b) => ({ x: (a.x + b.x) / 2, y: (a.y + b.y) / 2 }); +const dist = (a, b) => Math.hypot(a.x - b.x, a.y - b.y); + +// Which iris block belongs to which eye is RESOLVED FROM THE DATA, not declared +// in a table. +// +// The naming in MediaPipe's own material is viewer-relative in some places and +// subject-relative in others, and the two blocks are otherwise +// indistinguishable. Getting it backwards swaps the irises, which looks almost +// right - each eye still has a disc in roughly the right place - so it survives +// a casual eyeball and then reads as a subtly wall-eyed character for the rest +// of the project. Proximity to the eye's corner midpoint settles it in one +// comparison, is impossible to get wrong, and keeps working if the model is +// ever renumbered. +// +// Voted across every frame rather than read off frame zero: one bad detection +// should not decide the whole shot. +export function pairIrises(stab) { + if (!stab.irisA) return null; + let votes = 0; + for (let f = 0; f < stab.irisA.length; f++) { + const cR = mid(stab.cornersR[f][0], stab.cornersR[f][1]); + votes += dist(stab.irisA[f][0], cR) < dist(stab.irisB[f][0], cR) ? 1 : -1; + } + return votes > 0 ? { right: 'irisA', left: 'irisB' } + : { right: 'irisB', left: 'irisA' }; +} + +// Per-frame eye measurements, in units of eye width. Measurement only - every +// threshold and every stylisation is applied by the callers. +// +// Everything here stays in HEAD-LOCAL space, which is the same space the mouth +// lives in and the same space the registered photo underlay is drawn in. An +// earlier version pinned each eye into a fixed socket fitted to its corners' +// mean over the shot. That does remove the wobble, but it removes too much: the +// residual from out-of-plane rotation is real motion of the eye relative to the +// head, it is still there in the footage, and pinning it away leaves the drawn +// eyes hanging still over a photo whose eyes are moving. The eye has to track +// the face exactly as the mouth does. +// +// The wobble the socket was aimed at is dealt with the way docs/design.md deals +// with it everywhere else - the bounded contour average, the same knob and the +// same radius the mouth uses - and by placing the iris in the frame of the +// ALREADY-SMOOTHED lid ring, so the iris cannot jitter independently of the eye +// it sits in. See buildEyes in app.js. +export function eyeSignals(stab) { + const N = stab.transforms.length; + const pairing = pairIrises(stab); + const openR = [], openL = [], gazeRaw = [], gazeR = [], gazeL = []; + + for (let f = 0; f < N; f++) { + const cR = mid(stab.cornersR[f][0], stab.cornersR[f][1]); + const cL = mid(stab.cornersL[f][0], stab.cornersL[f][1]); + const wR = dist(stab.cornersR[f][0], stab.cornersR[f][1]); + const wL = dist(stab.cornersL[f][0], stab.cornersL[f][1]); + + // Openness is the lid gap over the CORNER distance. Normalising by the + // corners rather than by anything derived from the lids keeps the + // denominator rigid, so the ratio measures the lid and nothing else, and + // one threshold carries across takes, faces and framings. + openR.push(dist(stab.lidsR[f][0], stab.lidsR[f][1]) / wR); + openL.push(dist(stab.lidsL[f][0], stab.lidsL[f][1]) / wL); + + if (!pairing) { + gazeRaw.push({ x: 0, y: 0 }); gazeR.push({ x: 0, y: 0 }); gazeL.push({ x: 0, y: 0 }); + continue; + } + const iR = stab[pairing.right][f][0], iL = stab[pairing.left][f][0]; + // Gaze is the iris centre relative to the CORNER MIDPOINT, in eye widths - + // a pure offset WITHIN the eye, with the eye's own position divided out, so + // that quantising it quantises the glance and not the head motion carrying + // it. + // + // Measuring against the lid ring's centroid instead would track the lid: + // every blink pulls that centroid down and would fake a glance at the + // floor, on precisely the frames where the eye is most conspicuous. The + // corners are in RIGID, so this origin and this denominator are both immune + // to the performance they are measuring. + const gR = { x: (iR.x - cR.x) / wR, y: (iR.y - cR.y) / wR }; + const gL = { x: (iL.x - cL.x) / wL, y: (iL.y - cL.y) / wL }; + + // ONE gaze for both eyes, and deliberately so. At 320x200 an iris is a + // handful of pixels and its centre comes from five landmarks on an eye + // twenty pixels wide, so the difference between the two measurements is + // noise, not vergence - and independent per-eye noise reads as wall-eyed + // immediately, which is the most expensive artefact on a face. Openness + // stays per-eye, because a wink is real performance and should survive. + gazeRaw.push({ x: (gR.x + gL.x) / 2, y: (gR.y + gL.y) / 2 }); + // Kept separately purely as a diagnostic. The two eyes should agree; when + // they disagree in a sustained way rather than frame to frame, that is not + // noise but out-of-plane head rotation biasing the projected iris offset, + // and no 2D measurement can undo it. + gazeR.push(gR); gazeL.push(gL); + } + return { openR, openL, gazeRaw, gazeR, gazeL, hasIris: !!pairing }; +} + +// Where "not looking anywhere in particular" sits on THIS face. Everything the +// character does is measured as a departure from it, so getting it wrong does +// not bias the gaze slightly - it re-points the whole performance. +// +// `median` is the default and the safe one: the middle of the take, per axis. +// docs/design.md already gives this rule for the anchor fit - the reference is +// the MEAN configuration over the shot, not one frame - and gaze needs it for +// the same reason. The median rather than the mean because a couple of frames +// of hard glance should not drag the rest-point after them. +// +// `neutral` reads the origin off the take's neutral frame instead, which is +// only correct when there genuinely is a held neutral to read. That frame is +// chosen by MINIMUM MOUTH APERTURE, and a closed mouth says nothing whatever +// about where the eyes are pointed - so on footage with no deliberate neutral +// at the top it is an arbitrary frame, and whichever way the performer happened +// to glance on it becomes "straight ahead" for the entire shot. It is kept +// because it is right when the take was shot for this tool, and because being +// able to switch is how you find out that it was not. +export function gazeOrigin(gazeRaw, mode = 'median', neutral = 0, radius = 2) { + if (mode === 'neutral') { + let sx = 0, sy = 0, n = 0; + // A window, not a single frame: one frame of a five-landmark iris centre is + // worth about a pixel of noise, and that pixel would become a permanent + // squint in the output. + for (let f = neutral - radius; f <= neutral + radius; f++) { + const k = Math.min(gazeRaw.length - 1, Math.max(0, f)); + sx += gazeRaw[k].x; sy += gazeRaw[k].y; n++; + } + return { x: sx / n, y: sy / n }; + } + const mid1 = (vals) => { + const v = vals.slice().sort((a, b) => a - b); + return v.length % 2 ? v[(v.length - 1) / 2] + : (v[v.length / 2 - 1] + v[v.length / 2]) / 2; + }; + return { x: mid1(gazeRaw.map((g) => g.x)), y: mid1(gazeRaw.map((g) => g.y)) }; +} + +// Snap a two-channel track onto a grid, then require a new cell to hold before +// it takes. Gaze uses it for (x, y); brows use it for (outer raise, inner raise), +// where sharing the dwell is the point - a brow whose inner end arrived a frame +// before its outer end would crawl instead of snapping. +// +// This is the "Primitive - quantised" row of the part table in docs/design.md, +// and it is not a stylisation imposed on the truth: real eyes move in saccades, +// holding a fixation and then jumping. The smooth drift left in the measurement +// is tracker noise plus head-compensation error, so snapping to a grid and +// requiring a dwell removes the noise and recovers the saccade in the same +// operation - the rare case where the aesthetic rule and the physiology agree. +// +// The dwell is what stops a gaze parked on a cell boundary from chattering +// between two cells forever. It is meaningless without a grid, because +// continuous values never repeat, so step 0 short-circuits both. +export function quantizeSnap(track, step, dwell) { + if (!(step > 0)) return track.map((g) => ({ x: g.x, y: g.y })); + const q = track.map((g) => ({ + x: Math.round(g.x / step) * step, + y: Math.round(g.y / step) * step, + })); + if (dwell <= 0 || !q.length) return q; + + const out = []; + let live = q[0], pend = q[0], run = 0; + for (const g of q) { + if (g.x === pend.x && g.y === pend.y) run++; + else { pend = g; run = 1; } + if (run > dwell && (pend.x !== live.x || pend.y !== live.y)) live = pend; + out.push(live); + } + return out; +} + +// Resolve openness into a shut/open decision per frame. +// +// `dwell` is the same guard the teeth get: a lid hovering at the threshold must +// commit before the state changes, so it cannot flicker. +// +// `hold` is the one that is NOT like the teeth, and it is the whole reason +// blinks are worth special-casing. A blink is 100-150ms, which at 12fps is one +// frame and at 24fps is two or three - and a single frame of closed eye reads +// as a dropped frame, not as a blink. Animators draw a blink over two or three +// drawings for exactly that reason. So once the eye shuts it stays shut for +// `hold` frames, which turns an unreadable flicker into a beat. +// +// The hysteresis runs the other way from the teeth: shutting needs a clear +// signal, and once shut the eye is given the benefit of the doubt on reopening, +// because the lid landmarks are least reliable mid-blink. +export function resolveBlink(open, { cut, dwell, hold }) { + const N = open.length; + const shut = new Array(N).fill(false); + let live = false; // current state + let run = 0; // frames the opposing reading has persisted + let held = 0; // frames spent in the current state + for (let f = 0; f < N; f++) { + const reading = live ? open[f] < cut * 1.35 : open[f] < cut; + if (reading === live) run = 0; + else { + run++; + // Leaving a blink additionally requires the blink to have been on screen + // long enough to be legible; entering one never waits. + if (run > dwell && (!live || held >= hold)) { live = reading; held = 0; run = 0; } + } + held++; + shut[f] = live; + } + return shut; +} + // Stage 4: fixed-index subsample of a stabilised ring, then map from normalised // face space into character raster space. export function toRasterRing(stabRing, ringTable, n, xform) { @@ -178,6 +475,24 @@ export function heldFrame(kept, f) { return hit; } +// Hold every output frame back onto an exposure grid: 1 = on 1s, 2 = on 2s, and +// so on. Frame 5 at exposure 2 reads the pose from frame 4. +// +// This is where "aesthetic sparseness" belongs. docs/design.md used to put it at +// the extraction rate - pick 12fps and the timing is already chosen - but that +// makes the timing a property of a directory of PNGs, so auditioning 12 against +// 24 means re-ripping the clip and re-running detection over all of it. Rip +// dense once and quantise here instead: the dense track stays at the camera's +// rate, the decision stays reversible, and the audio clock is untouched, so +// sync cannot drift while you try timings. +// +// Floor, never round. Rounding would let an output frame read a pose from the +// FUTURE, which is a lead - a separate control, applied after this one, for a +// separate reason. +export function exposeIndex(f, exposure) { + return exposure > 1 ? Math.floor(f / exposure) * exposure : f; +} + // Shift a performance track against the clock, clamped at the ends. // // Pure and exported so the shift can actually be asserted: "the slider feels diff --git a/js/raster.js b/js/raster.js index 477e63f..987fbe8 100644 --- a/js/raster.js +++ b/js/raster.js @@ -46,14 +46,47 @@ export class IndexedRaster { } } - fillDisc(cx, cy, r, index) { + // `over` is an optional stencil: when given, only pixels that currently hold + // that index are written. The indexed buffer is its own clip mask, which is + // how Animator Pro would do it - and it is what keeps the iris inside the + // eye. A disc clipped by the sclera cannot spill past the lid at any gaze or + // any radius, including mid-blink when the opening is a two-pixel sliver, so + // the lid crops the iris for free instead of the gaze range needing a + // clamp that would flatten the performance at the extremes. + fillDisc(cx, cy, r, index, over = null) { const rr = r * r; const y0 = Math.max(0, Math.floor(cy - r)), y1 = Math.min(this.h - 1, Math.ceil(cy + r)); const x0 = Math.max(0, Math.floor(cx - r)), x1 = Math.min(this.w - 1, Math.ceil(cx + r)); for (let y = y0; y <= y1; y++) { for (let x = x0; x <= x1; x++) { const dx = x + 0.5 - cx, dy = y + 0.5 - cy; - if (dx * dx + dy * dy <= rr) this.buf[y * this.w + x] = index; + if (dx * dx + dy * dy > rr) continue; + const o = y * this.w + x; + if (over === null || this.buf[o] === over) this.buf[o] = index; + } + } + } + + // An exactly size x size block of pixels, snapped to the pixel grid, with the + // same optional stencil as fillDisc. + // + // The pupil is a SQUARE because at 320x200 it is three pixels across, and a + // circle of radius 1.5 is not a circle - it is a plus sign with the corners + // gnawed off, and it changes shape as it moves. A square that size is a + // deliberate mark that stays the same mark wherever it lands, which is the + // whole argument for flat shapes at this resolution. + // + // The top-left is rounded rather than the centre, so the block is size x size + // on every frame. Round the extents instead and a fractional centre gives you + // three pixels on one frame and four on the next, which reads as the pupil + // breathing. + fillRect(cx, cy, size, index, over = null) { + if (size < 1) return; + const x0 = Math.round(cx - size / 2), y0 = Math.round(cy - size / 2); + for (let y = Math.max(0, y0); y < Math.min(this.h, y0 + size); y++) { + for (let x = Math.max(0, x0); x < Math.min(this.w, x0 + size); x++) { + const o = y * this.w + x; + if (over === null || this.buf[o] === over) this.buf[o] = index; } } } diff --git a/js/selftest.js b/js/selftest.js index 026c476..3ddf456 100644 --- a/js/selftest.js +++ b/js/selftest.js @@ -7,9 +7,14 @@ // as blocks meeting at corners. It is invisible at some vertex counts and obvious // at others, so it needs an assertion rather than an eyeball. -import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, subsampleSlots, subsampleRing } from './landmarks.js'; -import { fitSimilarity, applySim, procrustesMean, smoothTransforms } from './mathutil.js'; -import { stabilize, toRasterRing, selectKeys, activeKey, shiftIndex } from './pipeline.js'; +import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, subsampleSlots, subsampleRing, + EYE_R_RING, EYE_L_RING, EYE_R_CORNERS, EYE_L_CORNERS, + EYE_R_LIDS, EYE_L_LIDS, BROW_A_RING, BROW_B_RING, + BROW_END_0, BROW_END_1 } from './landmarks.js'; +import { fitSimilarity, applySim, procrustesMean, smoothTransforms, offsetRing } from './mathutil.js'; +import { stabilize, toRasterRing, smoothContours, selectKeys, activeKey, shiftIndex, + exposeIndex, eyeSignals, pairIrises, gazeOrigin, quantizeSnap, resolveBlink, + browSignals, pairBrows } from './pipeline.js'; import { IndexedRaster, hexToRgb } from './raster.js'; import { writeTake } from './take.js'; import { otsuForTest, scaleRing } from './interior.js'; @@ -206,6 +211,25 @@ export function run() { ok('activeKey holds between keys', activeKey(sel.keys, sel.keys[1].f - 1).f === sel.keys[0].f); + // exposure: rip dense, choose the timing here. On 2s every odd frame must + // reuse the even frame's pose, and the grid must never read from the future - + // that direction is the lead, which is a different control for a reason. + ok('exposure 1 is identity', [0, 1, 7, 71].every((f) => exposeIndex(f, 1) === f)); + ok('on 2s holds each pose for two frames', + [0, 1, 2, 3, 4, 5].map((f) => exposeIndex(f, 2)).join(',') === '0,0,2,2,4,4'); + ok('on 3s holds each pose for three frames', + [0, 1, 2, 3, 4, 5, 6].map((f) => exposeIndex(f, 3)).join(',') === '0,0,0,3,3,3,6'); + ok('exposure never reads a pose from the future', + [0, 1, 2, 3, 4, 5, 6, 7].every((f) => exposeIndex(f, 3) <= f)); + { + // Exposure then lead, in that order: the picture must change on the grid + // beats and carry a pose shifted by whole frames of the original track. + const N = 72, at = (f) => shiftIndex(exposeIndex(f, 2), 1, N); + ok('exposure and lead compose without moving the beats', + at(0) === 1 && at(1) === 1 && at(2) === 3 && at(3) === 3, + [0, 1, 2, 3].map(at).join(',')); + } + // mouth lead: a shift that "feels like it does nothing" is indistinguishable // from one that does nothing, so assert the arithmetic directly. ok('lead 0 is identity', [0, 5, 71].every((f) => shiftIndex(f, 0, 72) === f)); @@ -288,6 +312,333 @@ export function run() { (tt.mBright - tt.mDark) / 255 > 0.4, `sep ${((tt.mBright - tt.mDark) / 255).toFixed(4)}`); } + /* ---- eyes ---- */ + + ok('eye rings have 16 distinct ids each', + new Set(EYE_R_RING).size === 16 && new Set(EYE_L_RING).size === 16); + ok('the two eye rings share no landmark', + !EYE_R_RING.some((i) => EYE_L_RING.includes(i))); + + // The cardinal contract, asserted rather than trusted: on a 16-slot ring the + // quarter slots must be the four anatomical cardinals, which is what makes + // every even vertex budget land on real landmarks instead of between them. + ok('eye ring slot 0/4/8/12 are outer, upper, inner, lower', + EYE_R_RING[0] === EYE_R_CORNERS[0] && EYE_R_RING[8] === EYE_R_CORNERS[1] && + EYE_R_RING[4] === EYE_R_LIDS[0] && EYE_R_RING[12] === EYE_R_LIDS[1] && + EYE_L_RING[0] === EYE_L_CORNERS[0] && EYE_L_RING[8] === EYE_L_CORNERS[1] && + EYE_L_RING[4] === EYE_L_LIDS[0] && EYE_L_RING[12] === EYE_L_LIDS[1]); + + // The gaze origin and denominator are built from the eye corners, so if a + // corner were not rigid a blink could move it and fake a glance. + ok('every eye corner is a rigid landmark', + [...EYE_R_CORNERS, ...EYE_L_CORNERS].every((i) => RIGID.includes(i))); + + // Same simplicity requirement as the lips, and for the same reason: a cut + // part with a self-intersecting ring renders as blocks meeting at corners. + // Checked on blink frames too, where the ring is nearly degenerate. + for (const [label, table] of [['right', EYE_R_RING], ['left', EYE_L_RING]]) { + let worst = null; + for (let n = 4; n <= 12 && !worst; n += 2) { + const slots = subsampleSlots(table.length, n); + for (let f = 0; f < dense.length; f++) { + const hits = ringSelfIntersections(slots.map((sl) => dense[f][table[sl]])); + if (hits.length) { worst = `verts=${n} frame=${f} edges ${JSON.stringify(hits[0])}`; break; } + } + } + ok(`${label} eye ring is simple at every vertex budget`, !worst, worst || ''); + } + + // offsetRing must grow by a FIXED amount and survive a degenerate ring - the + // shut eyelid is exactly the degenerate case, and it is the frame where the + // lash line is the entire drawing. + { + const sq = [{ x: -1, y: 0 }, { x: 0, y: -1 }, { x: 1, y: 0 }, { x: 0, y: 1 }]; + const g = offsetRing(sq, 2); + ok('offsetRing pushes every vertex out by exactly d', + g.every((p, i) => Math.abs(Math.hypot(p.x, p.y) - (Math.hypot(sq[i].x, sq[i].y) + 2)) < 1e-9)); + ok('offsetRing(0) is identity', offsetRing(sq, 0) === sq); + // A shut lid: a flat sliver. The offset must still open it into a band. + const shutLid = [{ x: -10, y: 0 }, { x: 0, y: -0.02 }, { x: 10, y: 0 }, { x: 0, y: 0.02 }]; + const band = offsetRing(shutLid, 1.5); + const h = Math.max(...band.map((p) => p.y)) - Math.min(...band.map((p) => p.y)); + ok('offsetRing gives a shut lid a visible lash band', h > 2.9, `height ${h.toFixed(3)}`); + ok('offsetRing keeps the shut lid simple', ringSelfIntersections(band).length === 0); + } + + { + const stE = stabilize(dense, 2); + const sig = eyeSignals(stE); + ok('synthetic track carries iris landmarks', sig.hasIris); + + // THE load-bearing eye assertion. The pairing is resolved from geometry + // rather than declared, so the test feeds a track built the OTHER way round + // and demands the resolver follow the data. A resolver only ever checked + // against the convention it was written for is checking nothing. + const pairA = pairIrises(stE); + const pairB = pairIrises(stabilize(synthDense(72, { swapIris: true }), 2)); + ok('iris pairing is resolved from the data, not assumed', + pairA.right === 'irisA' && pairB.right === 'irisB', + `normal ${pairA.right}, swapped ${pairB.right}`); + + // The eye must TRACK the face, not sit in a fixed socket. An earlier + // version pinned each eye to its corners' mean over the shot, which does + // kill the wobble but leaves the drawn eyes hanging still over a registered + // photo whose eyes are moving. Head-local is the same space the mouth and + // the underlay live in, so the eye moves with the head exactly as they do. + { + const spread = (arr, sel) => { + const v = arr.map(sel); + return Math.max(...v) - Math.min(...v); + }; + const w = Math.hypot(stE.cornersR[0][0].x - stE.cornersR[0][1].x, + stE.cornersR[0][0].y - stE.cornersR[0][1].y); + const moves = Math.max(spread(stE.lidR, (r) => r[0].x), spread(stE.lidR, (r) => r[0].y)); + ok('the eye stays in head-local space and tracks the face', moves / w > 0.02, + `corner travels ${(moves / w * 100).toFixed(1)}% of an eye width`); + + // Subsampling a 16-slot ring to any even budget must keep the two corners + // at output indices 0 and n/2. That is what lets the socket be read back + // off the drawn polygon instead of measured separately, which is what + // stops the iris drifting relative to the eye it sits in. + let bad = null; + for (let n = 4; n <= 12; n += 2) { + const sl = subsampleSlots(16, n); + if (sl[0] !== 0 || sl[n / 2] !== 8) bad = `n=${n} -> ${sl.join(',')}`; + } + ok('the drawn lid ring carries its own corners at 0 and n/2', !bad, bad || ''); + + // The contour average is what removes the jitter, and it is the mouth's + // knob doing the mouth's job - no second mechanism for the eyes. + const ring = (rad) => smoothContours( + stE.lidR.map((r) => toRasterRing(r, EYE_R_RING, 8, (p) => ({ x: p.x * 600, y: p.y * 600 }))), rad); + const jitter = (rings) => { + let acc = 0; + for (let f = 1; f < rings.length; f++) { + const a = rings[f], b = rings[f - 1]; + acc += Math.hypot((a[0].x + a[4].x) / 2 - (b[0].x + b[4].x) / 2, + (a[0].y + a[4].y) / 2 - (b[0].y + b[4].y) / 2); + } + return acc / (rings.length - 1); + }; + ok('contour averaging steadies the eye without pinning it', + jitter(ring(1)) < jitter(ring(0)) * 0.8, + `${jitter(ring(0)).toFixed(3)} -> ${jitter(ring(1)).toFixed(3)} px/frame`); + } + + // Blink: synth shuts the lids for exactly one frame every 19. + const lo = Math.min(...sig.openR), hi = Math.max(...sig.openR); + ok('openness collapses on a blink and not otherwise', lo < hi * 0.2, + `${lo.toFixed(3)} .. ${hi.toFixed(3)}`); + + const shut = resolveBlink(sig.openR, { cut: hi * 0.3, dwell: 0, hold: 3 }); + const runs = []; + for (let f = 0; f < shut.length; f++) if (shut[f] && !shut[f - 1]) runs.push(f); + const lens = runs.map((a) => { let n = 0; while (shut[a + n]) n++; return n; }); + ok('blinks are found', runs.length >= 3, `${runs.length} runs at ${runs.join(',')}`); + // The knob that is not like the teeth: a one-frame blink reads as a dropped + // frame, so `hold` must stretch it into something legible. + ok('a one-frame blink is held to the minimum length', + lens.every((n) => n >= 3), `run lengths ${lens.join(',')}`); + ok('a shorter hold leaves the blink shorter', + resolveBlink(sig.openR, { cut: hi * 0.3, dwell: 0, hold: 1 }).filter(Boolean).length < + shut.filter(Boolean).length); + + // Gaze, against ground truth: synth commands +0.16 eye widths at f12 and + // -0.16 at f23, holding each for eleven frames. + const org = gazeOrigin(sig.gazeRaw, 'neutral', 0); + const gx = (f) => (sig.gazeRaw[f].x - org.x); + ok('gaze recovers the commanded direction', + gx(12) > 0.12 && gx(12) < 0.20 && gx(23) < -0.12 && gx(23) > -0.20, + `f12 ${gx(12).toFixed(3)}, f23 ${gx(23).toFixed(3)}`); + + // Measuring gaze against the lid centroid instead of the corner midpoint + // would drag the iris down on every blink and fake a glance at the floor, + // on exactly the frames where the eye is most conspicuous. + const gy = (f) => (sig.gazeRaw[f].y - org.y); + ok('a blink does not fake a change of gaze', + Math.abs(gy(19) - gy(18)) < 0.02, `f18 ${gy(18).toFixed(4)} -> f19 ${gy(19).toFixed(4)}`); + + // Quantisation is what turns drift into saccades: four commanded + // fixations must come back as a handful of cells, not one per frame. + // The origin re-points the whole performance, so a wrong one does not bias + // the gaze slightly - it makes the character look the other way. The median + // must sit inside the range it summarises; the neutral-frame origin need + // not, which is exactly the failure mode it has on footage with no + // deliberate neutral at the top. + { + const med = gazeOrigin(sig.gazeRaw, 'median'); + const xs = sig.gazeRaw.map((g) => g.x); + ok('the median origin lies inside the take\'s own gaze range', + med.x > Math.min(...xs) && med.x < Math.max(...xs), + `${med.x.toFixed(3)} in ${Math.min(...xs).toFixed(3)}..${Math.max(...xs).toFixed(3)}`); + // Synth looks left as much as right, so the rest point is near zero. + ok('the median origin finds the rest point, not a glance', + Math.abs(med.x) < 0.08, `median x ${med.x.toFixed(3)}`); + ok('the two origins actually differ, so the toggle is a real A/B', + Math.abs(med.x - gazeOrigin(sig.gazeRaw, 'neutral', 12).x) > 0.02); + } + + const px = sig.gazeRaw.map((g) => ({ x: (g.x - org.x) * 30, y: (g.y - org.y) * 30 })); + const cells = (a) => new Set(a.map((g) => `${g.x},${g.y}`)).size; + ok('quantisation collapses drift into a few fixations', + cells(quantizeSnap(px, 2, 2)) <= 6 && cells(px) > 40, + `${cells(px)} raw -> ${cells(quantizeSnap(px, 2, 2))} cells`); + ok('gaze step 0 leaves the track untouched', + quantizeSnap(px, 0, 2).every((g, i) => g.x === px[i].x && g.y === px[i].y)); + ok('quantised values land on the grid', + quantizeSnap(px, 2, 0).every((g) => Math.abs(g.x % 2) < 1e-9 && Math.abs(g.y % 2) < 1e-9)); + // A one-frame excursion is noise; the dwell must swallow it. + { + const spike = [{ x: 0, y: 0 }, { x: 0, y: 0 }, { x: 4, y: 0 }, { x: 0, y: 0 }, { x: 0, y: 0 }]; + ok('the dwell suppresses a one-frame gaze spike', + quantizeSnap(spike, 2, 1).every((g) => g.x === 0)); + ok('a sustained move still gets through', + quantizeSnap([...spike, { x: 4, y: 0 }, { x: 4, y: 0 }, { x: 4, y: 0 }], 2, 1).pop().x === 4); + } + } + + /* ---- brows ---- */ + + ok('brow rings have 10 distinct ids each', + new Set(BROW_A_RING).size === 10 && new Set(BROW_B_RING).size === 10); + ok('the brow rings share no landmark with each other, RIGID, or the lids', + !BROW_A_RING.some((i) => BROW_B_RING.includes(i)) && + ![...BROW_A_RING, ...BROW_B_RING].some((i) => + RIGID.includes(i) || EYE_R_RING.includes(i) || EYE_L_RING.includes(i)), + 'a brow in RIGID would bleed expression into the stabilisation'); + + { + let bad = null; + for (const [label, table] of [['A', BROW_A_RING], ['B', BROW_B_RING]]) { + for (let n = 4; n <= 10 && !bad; n += 2) { + const slots = subsampleSlots(table.length, n); + for (let f = 0; f < dense.length; f++) { + if (ringSelfIntersections(slots.map((sl) => dense[f][table[sl]])).length) { + bad = `${label} verts=${n} frame=${f}`; break; + } + } + } + } + ok('brow rings are simple at every vertex budget', !bad, bad || ''); + } + + { + const stB = stabilize(dense, 2); + const pr = pairBrows(stB); + ok('brow-to-eye pairing is resolved from geometry', + pr.right === 'browA' && pr.left === 'browB', JSON.stringify(pr)); + // Getting this backwards mirrors the tilt, so inner-up "worried" renders as + // outer-up. That is a different expression, not a broken one, which is + // exactly why it needs an assertion rather than an eyeball. + ok('the outer end of the brow ring is resolved from geometry', pr.outerAtSlot0 === true); + + // Both ends land on fixed slots whichever edge of the brow is on top, which + // is what lets the upper/lower ambiguity go unresolved without consequence. + ok('brow end slots are disjoint and cover both ends', + !BROW_END_0.some((i) => BROW_END_1.includes(i)) && + BROW_END_0.length === 2 && BROW_END_1.length === 2); + + // Ground truth: synth commands rest, surprise, worry and anger as heights + // above the eye centre in eye widths, holding each for thirteen frames. + const b = browSignals(stB); + const at = (f) => [b.R[f].x, b.R[f].y]; + const near = (v, want) => Math.abs(v - want) < 0.02; + ok('brow raise recovers the commanded rest pose', near(at(0)[0], 0.30) && near(at(0)[1], 0.30), + at(0).map((v) => v.toFixed(3)).join(', ')); + ok('brow raise recovers surprise - both ends up', + near(at(14)[0], 0.46) && near(at(14)[1], 0.46), at(14).map((v) => v.toFixed(3)).join(', ')); + ok('brow raise recovers worry - inner end only', + near(at(27)[0], 0.30) && near(at(27)[1], 0.44), at(27).map((v) => v.toFixed(3)).join(', ')); + ok('brow raise recovers anger - inner end down', + near(at(40)[0], 0.30) && near(at(40)[1], 0.18), at(40).map((v) => v.toFixed(3)).join(', ')); + + // Tilt must be a signed quantity that separates worry from anger. If the + // outer/inner resolution were mirrored these two would swap. + ok('tilt separates worry from anger by sign', + (at(27)[1] - at(27)[0]) > 0.08 && (at(40)[1] - at(40)[0]) < -0.08, + `worry ${(at(27)[1] - at(27)[0]).toFixed(3)}, anger ${(at(40)[1] - at(40)[0]).toFixed(3)}`); + + // A blink must not read as a brow raise: the raise is measured against the + // eye's rigid corners, not its lid, which is the same trap the gaze origin + // has and worth avoiding twice. + const dR = Math.abs(b.R[19].x - b.R[18].x); + ok('a blink does not fake a brow raise', dR < 0.01, `f18 -> f19 delta ${dR.toFixed(4)}`); + + // Quantisation: four sustained poses must come back as a handful of levels. + const rest = gazeOrigin(b.R, 'median'); + const px = b.R.map((g) => ({ x: (g.x - rest.x) * 60, y: (g.y - rest.y) * 60 })); + const cells = (a) => new Set(a.map((g) => `${g.x},${g.y}`)).size; + ok('brow quantisation collapses drift into a few poses', + cells(quantizeSnap(px, 2, 2)) <= 6 && cells(px) > 20, + `${cells(px)} raw -> ${cells(quantizeSnap(px, 2, 2))} poses`); + // The dwell is SHARED across both channels, and that is the whole reason + // brows reuse the gaze quantiser rather than running two independent ones. + // Here the outer end moves one frame before the inner: with a shared dwell + // the half-raised pose (2,0) is transient and never commits, so the brow + // snaps once. Two independent dwells would emit it and the brow would crawl + // into position over two frames instead of hitting it. + { + const staggered = [ + { x: 0, y: 0 }, { x: 0, y: 0 }, { x: 2, y: 0 }, + { x: 2, y: 2 }, { x: 2, y: 2 }, { x: 2, y: 2 }, + ]; + const out = quantizeSnap(staggered, 2, 1); + ok('a shared dwell never emits a half-raised brow', + !out.some((g) => g.x === 2 && g.y === 0), + out.map((g) => `${g.x},${g.y}`).join(' ')); + } + } + + // The iris is stencilled by the sclera and the pupil by the iris, which is + // what keeps both inside the lid at any gaze without clamping the gaze itself. + { + const rr = new IndexedRaster(40, 40); + rr.clear(0); + rr.fillPoly([{ x: 10, y: 10 }, { x: 30, y: 10 }, { x: 30, y: 20 }, { x: 10, y: 20 }], 1); + rr.fillDisc(28, 15, 9, 2, 1); // a disc reaching well past the "lid" + let spill = 0, inside = 0; + for (let y = 0; y < 40; y++) for (let x = 0; x < 40; x++) { + const v = rr.buf[y * 40 + x]; + if (v !== 2) continue; + if (x >= 10 && x < 30 && y >= 10 && y < 20) inside++; else spill++; + } + ok('a stencilled disc cannot spill past its clip', spill === 0 && inside > 20, + `${inside} in, ${spill} out`); + rr.fillDisc(5, 35, 3, 3); // no stencil: writes freely + ok('an unstencilled disc still writes anywhere', rr.buf.includes(3)); + + // The stencil chain: pupil over iris over sclera. A pupil placed where the + // iris has already been cropped must be cropped the same way. + rr.fillRect(28, 15, 5, 4, 2); + let pSpill = 0; + for (let y = 0; y < 40; y++) for (let x = 0; x < 40; x++) { + if (rr.buf[y * 40 + x] === 4 && !(x >= 10 && x < 30 && y >= 10 && y < 20)) pSpill++; + } + ok('the pupil inherits the iris clip transitively', pSpill === 0); + } + + // A square pupil is only worth having if it is the SAME square every frame: + // exactly its nominal size at any centre, or it breathes as the gaze moves. + { + const sizes = []; + for (const [cx, cy] of [[20, 20], [20.5, 20.5], [20.49, 19.51], [21, 20]]) { + const rr = new IndexedRaster(40, 40); + rr.clear(0); + rr.fillRect(cx, cy, 3, 1); + let n = 0, minX = 99, maxX = -1, minY = 99, maxY = -1; + for (let y = 0; y < 40; y++) for (let x = 0; x < 40; x++) { + if (rr.buf[y * 40 + x] !== 1) continue; + n++; minX = Math.min(minX, x); maxX = Math.max(maxX, x); + minY = Math.min(minY, y); maxY = Math.max(maxY, y); + } + sizes.push(`${maxX - minX + 1}x${maxY - minY + 1}:${n}`); + } + ok('a 3px pupil is 3x3 at every centre', sizes.every((v) => v === '3x3:9'), sizes.join(' ')); + const rr = new IndexedRaster(40, 40); + rr.clear(0); rr.fillRect(20, 20, 0, 1); + ok('pupil size 0 draws nothing', !rr.buf.includes(1)); + } + // take writer round-trip const take = { name: 'test', frames: 72, width: 320, height: 200, exposure: 2, @@ -297,6 +648,12 @@ export function run() { { name: 'head', kind: 'plate', z: 0, interp: 'hold', keys: [{ f: 0, plate: 0 }] }, { name: 'mouth', kind: 'poly', z: 30, color: 'skin', interp: 'hold', keys: sel.keys.map((k) => ({ f: k.f, src: k.src, pts: shapes[k.src] })) }, + { name: 'iris_r', kind: 'disc', z: 22, color: 'iris', interp: 'hold', + parent: 'eye_r_in', clip: 'eye_r_in', + keys: [{ f: 0, src: 0, c: { x: 120.4, y: 88.7 }, r: 5.5 }, { f: 2, hidden: true }] }, + { name: 'pupil_r', kind: 'rect', z: 23, color: 'pupil', interp: 'hold', + parent: 'iris_r', clip: 'iris_r', + keys: [{ f: 0, src: 0, c: { x: 120, y: 89 }, size: 3 }] }, ], }; const text = writeTake(take); @@ -309,6 +666,13 @@ export function run() { }), `${keyLines.length} key lines`); ok('take declares a plate and a part table', /^plate\s+0/m.test(text) && /^part\s+mouth/m.test(text)); + ok('a disc part declares its clip', /^part\s+iris_r.*clip=eye_r_in/m.test(text)); + ok('a disc key is three integers', /^key\s+iris_r\s+f=0\s+src=0\s+disc=120,89,6$/m.test(text), + (text.split('\n').find((l) => l.startsWith('key iris_r')) || '').trim()); + ok('a hidden disc key emits hidden', /^key\s+iris_r\s+f=2\s+hidden$/m.test(text)); + ok('a pupil key is a square, not a tessellated polygon', + /^key\s+pupil_r\s+f=0\s+src=0\s+rect=120,89,3$/m.test(text) && + /^part\s+pupil_r.*clip=iris_r/m.test(text)); ok('coordinates are integers', !/-?\d+\.\d/.test(text.split('\n').filter((l) => l.startsWith('key')).join(''))); return results; diff --git a/js/synth.js b/js/synth.js index bf736be..84252ba 100644 --- a/js/synth.js +++ b/js/synth.js @@ -5,11 +5,17 @@ // verified without a video file. A synthetic face is also the only way to test // stabilisation against a KNOWN head motion, since real footage gives no ground // truth to compare against. -import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, EYE_INNER } from './landmarks.js'; +import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, + EYE_R_RING, EYE_L_RING, IRIS_A, IRIS_B, + BROW_A_RING, BROW_B_RING } from './landmarks.js'; const NUM = 478; -export function synthDense(nFrames = 72) { +// `swapIris` places the two iris blocks on the opposite eyes. It exists so the +// pairing resolver can be tested against a track it actually disagrees with: +// a resolver checked only against the convention it was written for is checking +// nothing at all. +export function synthDense(nFrames = 72, { swapIris = false } = {}) { const frames = []; for (let t = 0; t < nFrames; t++) { const pts = new Array(NUM); @@ -35,11 +41,70 @@ export function synthDense(nFrames = 72) { const openAmt = target[beat]; const wide = 0.10 + (beat === 1 ? 0.012 : beat === 3 ? -0.008 : 0); - place(RIGID[0], -0.075, -0.045); place(RIGID[1], -0.028, -0.043); - place(RIGID[2], 0.028, -0.043); place(RIGID[3], 0.075, -0.045); place(RIGID[4], 0.000, -0.050); place(RIGID[5], 0.000, -0.020); place(RIGID[6], 0.000, 0.012); - place(EYE_INNER[0], -0.028, -0.043); place(EYE_INNER[1], 0.028, -0.043); + + // Eyes. The corners (RIGID[0..3]) are placed BY the lid rings rather than + // separately, because they are slots 0 and 8 of those rings: writing them + // twice is how the mouth grew a bowtie, and a corner that disagrees with + // its own ring would make the eye self-intersect at some vertex budgets + // and not others. + // + // A blink is ONE frame, which is the honest hard case: at 12fps that is + // what a real blink costs, and it is exactly the length that reads as a + // dropped frame rather than as a blink unless `hold` extends it. + const blink = t > 5 && t % 19 === 0; + const openness = blink ? 0.05 : 1; + + // Gaze holds and then jumps, the way gaze actually behaves, with a little + // jitter on top so quantisation has noise to remove and the dwell has + // something to suppress. + const LOOK = [[0, 0], [0.16, 0.0], [-0.16, 0.05], [0.0, -0.09]]; + const [gx, gy] = LOOK[Math.floor(t / 11) % LOOK.length]; + const jit = () => (Math.random() - 0.5) * 0.012; + + // Half the corner separation, and the lid half-height at full open. + const EYE_RX = 0.0235, EYE_RY = 0.011, EYE_Y = -0.044; + const eye = (ring, cx, dir, iris) => { + const n = ring.length; + for (let k = 0; k < n; k++) { + // dir flips the traversal so each ring runs the direction its real + // table does: slot 0 outer corner, 4 upper lid, 8 inner, 12 lower. + const a = dir > 0 ? Math.PI + (k / n) * Math.PI * 2 : -(k / n) * Math.PI * 2; + place(ring[k], cx + EYE_RX * Math.cos(a), + EYE_Y + EYE_RY * openness * Math.sin(a)); + } + // Iris: centre first, then four ring points, as the refined mesh emits. + const ix = cx + (gx + jit()) * EYE_RX * 2, iy = EYE_Y + (gy + jit()) * EYE_RX * 2; + place(iris[0], ix, iy); + for (let k = 1; k < iris.length; k++) { + const a = ((k - 1) / (iris.length - 1)) * Math.PI * 2; + place(iris[k], ix + 0.008 * Math.cos(a), iy + 0.008 * Math.sin(a)); + } + }; + eye(EYE_R_RING, -0.0515, 1, swapIris ? IRIS_B : IRIS_A); + eye(EYE_L_RING, 0.0515, -1, swapIris ? IRIS_A : IRIS_B); + + // Brows, held in four sustained poses so raise quantisation has genuine + // plateaux to find: rest, surprise (both ends up), worry (inner up only), + // anger (inner down). Commanded in eye widths above the eye centre so the + // measurement can be checked against a number rather than an eyeball. + const BROW = [[0.30, 0.30], [0.46, 0.46], [0.30, 0.44], [0.30, 0.18]]; + const [bOut, bIn] = BROW[Math.floor(t / 13) % BROW.length]; + const EYE_W = EYE_RX * 2, HALF = 0.006; // ring half-thickness + const brow = (ring, cx, outerSign) => { + // Slots 0-4 are one edge outer->inner, 5-9 the other inner->outer, so the + // ends land on {0,9} and {4,5} exactly as the table promises. + const n = ring.length, half = n / 2; + for (let k = 0; k < n; k++) { + const along = k < half ? k / (half - 1) : (n - 1 - k) / (half - 1); + const rise = bOut + (bIn - bOut) * along; + place(ring[k], cx + outerSign * (EYE_RX - along * EYE_W) * 1.05, + EYE_Y - rise * EYE_W + (k < half ? -HALF : HALF)); + } + }; + brow(BROW_A_RING, -0.0515, -1); + brow(BROW_B_RING, 0.0515, 1); // Lip rings as ellipse arcs, traversed so ring ORDER matches the tables: // slot 0 = right corner, 5 = top centre, 10 = left corner, 15 = bottom diff --git a/js/take.js b/js/take.js index 8a87b09..97929f3 100644 --- a/js/take.js +++ b/js/take.js @@ -15,7 +15,12 @@ export function writeTake(take) { // v1 emits a single frozen plate derived from the face oval. A real project // replaces this with hand-drawn angles referenced by cel frame; the record // shape is the same either way. - L.push(`plate 0 kind=poly slot_mouth=${r(take.slot.x)},${r(take.slot.y)} scale=1.00 rot=0 squash=1.00`); + L.push(`plate 0 kind=poly slot_mouth=${r(take.slot.x)},${r(take.slot.y)}` + + (take.eyeSlots + ? ` slot_eye_r=${r(take.eyeSlots.r.cx)},${r(take.eyeSlots.r.cy)},${r(take.eyeSlots.r.w)}` + + ` slot_eye_l=${r(take.eyeSlots.l.cx)},${r(take.eyeSlots.l.cy)},${r(take.eyeSlots.l.w)}` + : '') + + ` scale=1.00 rot=0 squash=1.00`); L.push(''); for (const part of take.parts) { const bits = [`part ${part.name.padEnd(9)} kind=${part.kind} z=${part.z}`]; @@ -23,6 +28,10 @@ export function writeTake(take) { if (part.kind === 'poly') bits.push('closed=1 fill=1'); bits.push(`interp=${part.interp}`); if (part.parent) bits.push(`parent=${part.parent}`); + // `clip` names a part this one is stencilled by, not merely drawn after. + // The iris needs it: at an extreme gaze the disc reaches past the lid, and + // ordering alone would put it on the cheek. + if (part.clip) bits.push(`clip=${part.clip}`); L.push(bits.join(' ')); } L.push(''); @@ -30,6 +39,20 @@ export function writeTake(take) { for (const k of part.keys) { if (k.hidden) { L.push(`key ${part.name.padEnd(9)} f=${k.f} hidden`); continue; } if (part.kind === 'plate') { L.push(`key ${part.name.padEnd(9)} f=${k.f} plate=${k.plate}`); continue; } + // A disc is three numbers, so it gets its own key shape rather than being + // pre-tessellated into a polygon here: the renderer draws a real circle + // with hard edges, and a five-pixel iris approximated by a polygon would + // lose a pixel off its silhouette on some frames and not others. + // A square, in whole pixels, at an integer centre. Same argument as the + // disc: the renderer is told the shape, not a polygon approximating it. + if (part.kind === 'rect') { + L.push(`key ${part.name.padEnd(9)} f=${k.f} src=${k.src} rect=${r(k.c.x)},${r(k.c.y)},${k.size}`); + continue; + } + if (part.kind === 'disc') { + L.push(`key ${part.name.padEnd(9)} f=${k.f} src=${k.src} disc=${r(k.c.x)},${r(k.c.y)},${r(k.r)}`); + continue; + } const pts = k.pts.map((p) => `${r(p.x)},${r(p.y)}`).join(' '); // f is authoritative (what renders); src is the pre-snap extreme frame, // kept only as a tuning signal - a key dragged far means the minimum-hold diff --git a/manifest.json b/manifest.json index 9bf7b85..b67b3e5 100644 --- a/manifest.json +++ b/manifest.json @@ -1 +1 @@ -{"fps":24,"frames":74,"dir":"frames","audio":"audio.wav","source":"IMG_8486.MOV"} +{"fps":24,"frames":105,"dir":"frames","audio":"audio.wav","source":"ScreenRecording_09-24-2026 16-26-57_1.mov"} diff --git a/serve.py b/serve.py new file mode 100755 index 0000000..36139d3 --- /dev/null +++ b/serve.py @@ -0,0 +1,28 @@ +#!/usr/bin/env python3 +"""Dev server that refuses to cache anything. + +`python3 -m http.server` sends Last-Modified, and browsers cache ES modules on +it hard enough that a reload can serve a stale js/app.js against a fresh +index.html. That failure is silent and points nowhere near its cause: the new +knobs appear in the markup, nothing wires them, no error is raised, and the +symptom reads as "the feature you just added does not work". +""" +import sys +from http.server import SimpleHTTPRequestHandler, ThreadingHTTPServer + + +class NoCache(SimpleHTTPRequestHandler): + def end_headers(self): + self.send_header('Cache-Control', 'no-store, must-revalidate') + self.send_header('Expires', '0') + super().end_headers() + + def log_message(self, fmt, *args): + if '304' not in fmt % args: + super().log_message(fmt, *args) + + +if __name__ == '__main__': + port = int(sys.argv[1]) if len(sys.argv) > 1 else 8777 + print(f'http://127.0.0.1:{port} (no-store)') + ThreadingHTTPServer(('127.0.0.1', port), NoCache).serve_forever()