MediaPipe has no landmarks inside the lips, so teeth have to come from pixels.
Tracing the bright blob would give a new contour every frame with no vertex
correspondence - the exact boil docs/roto-puppet.md warns about.
So the measurement yields a scalar, not a shape: Otsu within the cavity, scanned
from the top for where the bright run stops, giving one line height per frame.
The teeth polygon is the inner lip ring clipped to that line, so the silhouette
is always the mouth's own shape and cannot disagree with the lips around it,
and the only per-frame variable is a single number that smooths trivially.
Presence gets hysteresis and minimum dwell, as plate selection does: a teeth
block blinking on and off for single frames is worse than one simply absent.
Tongue is not implemented. The same scalar approach would apply, gated on
redness rather than brightness, but it is not visible in the test footage - the
cavity reads dark with a bright upper-teeth band and nothing else.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The generated face oval was never going to be good enough to draw from:
MediaPipe's face oval is the FACE boundary, cut at the hairline and excluding
hair, ears, jaw and neck, so it is an egg by construction. Segmentation would
give a real head outline but costs a 16MB model and per-frame inference for a
shape that gets replaced by a drawing anyway.
So the plate layer becomes switchable, and the useful modes are photographic:
the source frame mapped into raster space through the same transform chain the
contours go through. Registration is the whole point - the head sits still and
a drawing traced from the underlay is already aligned to the mouth. An
unregistered underlay would be decoration.
- underlay.js: pixel->raster affine (a general affine, since MediaPipe
normalises x by width and y by height), registered draw, palette posterise
- plate modes: photo / photo dim / posterized / oval / oval+photo / none, B cycles
- worksheet cells are registered composites rather than raw crops
- Save frame 4x writes a 1280x800 PNG to draw on
- selftest: FACE_OVAL simplicity, which was never asserted; a wrong ordering
there reads as a lumpy plate rather than an obvious bowtie
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Lip sync cannot be judged silently. extract.sh now pulls the audio track and
writes manifest.json alongside the frames; the page reads the true extraction
rate from it rather than assuming one, since a guessed fps desynchronises
picture from sound - the one thing this view exists to show.
Audio is the clock: frame = floor(currentTime * fps). A slow render loop drops
frames instead of drifting, and half/quarter speed work via playbackRate with
the picture following for free. Scrubbing, stepping and clicking a thumbnail
all seek the audio too.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sparseness was being applied for two different reasons at once. Aesthetic
sparseness is set by the extraction rate; labour sparseness only binds on the
plate, because a human draws each one. The mouth is traced and therefore free,
and in limited animation lip sync is routinely the densest element - on 1s
while heads hold on 2s and 3s.
So: the mouth gets a key on every frame, and the frame strip is now the
editing surface for deciding which frames need their own plate drawing. All
frames start kept; delete the ones you don't want.
- strip of face-cropped thumbnails, keep/drop per frame, keyboard driven
- worksheet panel lists the drawings needed and the range each one holds
- Suggest runs error-tolerance decimation on head pose as a starting point
- export writes sparse plate keys + dense mouth keys, with a hold manifest
- smoothContours: bounded exception to "never smooth the contour", which held
only while keys were sparse enough to reject detector noise by sampling
- averages are now a RADIUS in frames: 0 is off, 1 is +-1
- GPU delegate falls back to CPU instead of failing
- #synth / #frames autorun for headless smoke tests
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Analysis half of the pipeline in docs/roto-puppet.md. Stabilises a face out
of a clip via a similarity fit on rigid landmarks, reduces the lip contour to
a fixed vertex budget, selects sparse keys on velocity minima, and previews
the result as flat indexed fills so timing can be judged without an Animator
Pro render.
- landmarks.js ordered lip/oval rings; slot position is vertex identity
- mathutil.js closed-form 2D similarity, Procrustes mean, transform smoothing
- pipeline.js stabilise -> subsample -> key-select
- raster.js indexed scanline fill, no antialiasing
- take.js take-file writer
- selftest.js 29 assertions, incl. ring simplicity at every vertex budget
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>