2026-09-24 14:38:07 -04:00
# roto
Video → **take file** builder for the Animator Pro rotoscope/puppet pipeline
described in `../docs/roto-puppet.md` .
This is the analysis and tuning half. It stabilises a face out of a clip, reduces
the lip contour to a handful of vertices, selects sparse keys on motion extremes,
and previews the result as flat indexed fills — so the *look and timing* can be
judged in seconds rather than through a minutes-long Animator Pro render. It
emits a `.take` file; nothing here touches Animator Pro.
All policy lives here. The take arrives at the renderer with its keys already
chosen.
## Run
```sh
python3 -m http.server 8777 # from this directory
# open http://127.0.0.1:8777
```
**Synthetic take** needs no video and exercises everything below detection.
For real footage:
```sh
2026-09-24 14:53:09 -04:00
./extract.sh /path/to/clip.mov 12 # -> frames/*.png, audio.wav, manifest.json
2026-09-24 14:38:07 -04:00
```
then **Load frames** . MediaPipe's wasm is fetched from jsdelivr on first use;
`face_landmarker.task` is local.
Frames are pre-extracted rather than decoded in the page because browser video
seeking is approximate and `requestVideoFrameCallback` only delivers frames at
playback speed — neither gives a deterministic per-frame pass.
2026-09-24 14:53:09 -04:00
`manifest.json` records the true extraction rate. The page reads it rather than
assuming, because a guessed fps desynchronises audio from picture — and sync is
the one thing this view exists to show.
**Audio is the playback clock**: `frame = floor(audio.currentTime * fps)` . A slow
render loop therefore drops frames instead of drifting, and ½x / ¼x work by
setting `playbackRate` with the picture following for free.
2026-09-24 14:38:07 -04:00
## Shooting for it
Near-frontal, good light, consistent scale, head reasonably still. Hold a neutral
closed mouth for a second at the top of the take: that frame is picked
automatically as the neutral and drives calibration and the placeholder plate.
Out-of-plane head rotation cannot be stabilised away by a 2D similarity
transform — the `residual` readout rises when it happens. The pipeline's answer
is hand-drawn head plates, which this tool does not yet do.
## Knobs
| Knob | What it does |
| --- | --- |
| vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. |
2026-09-24 15:34:08 -04:00
| mouth lead ±f | Shifts the performance tracks earlier against the audio and the head. `[` `]` . |
2026-09-24 14:53:09 -04:00
| contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. |
| anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform* . |
2026-09-24 14:38:07 -04:00
| closed-mouth cut | Aperture below which the mouth interior is emitted as `hidden` . |
2026-09-24 14:53:09 -04:00
| suggest tolerance | Max head movement before a new plate drawing is required. Affects **Suggest** only. |
2026-09-24 15:12:09 -04:00
## Teeth
MediaPipe has no landmarks inside the lips: the inner ring bounds the cavity and
Teeth as an extracted blob contour, not a clipped band
The band filled the mouth because a band is the wrong reduction: the bright
region is a blob, and reading it as "everything above a line" throws the shape
away.
Extracting a contour reintroduces the vertex-correspondence problem that made
me avoid it, but for a blob there is a way out. Radial sampling from the
centroid along N fixed directions makes vertex k always mean "the extent in
direction k": correspondence holds by construction, the count is fixed, and
temporal smoothing cannot reorder anything. It also yields a star-shaped
reduction, which suits flat colour.
Tongue rejection, which the band had no way to express:
- pixels red relative to their own brightness are dropped (teeth are neutral)
- component choice is biased toward the top of the cavity, since area alone
picks the tongue when the mouth is wide
- separate inner and outer controls: cavity erode pulls the sampled region off
the lip edge, blob grow/erode resizes the found blob
Also: a knob wired in app.js but missing from index.html threw during wiring
and left a blank page with nothing useful in the console - which is exactly
what happened to teethDwell in the previous commit. el() now names the missing
id, and window.onerror surfaces it in the status line.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:24:52 -04:00
everything within it is just pixels. So teeth come from the image.
The hazard is vertex correspondence — a traced contour reorders between frames
and boils. The way out for a blob specifically is **radial sampling** : march
outward from the blob's centroid along N fixed directions and take the last pixel
inside. Vertex *k* is then always "the extent in direction *k* ", so correspondence
holds by construction, the vertex count is fixed, and temporal smoothing is well
defined with no reordering possible. It also produces a star-shaped reduction,
which is what flat blocks of colour want.
The **teeth measurement** panel shows exactly what is sampled: region dimmed,
kept pixels green, extracted contour amber. Tune against that, not the numbers.
| Knob | What it does |
| --- | --- |
| teeth contrast | Gate on the separation between the cavity's dark and bright class means. Otsu always returns *some* threshold, so this is what stops it inventing teeth in a dark mouth. |
| cavity erode | Pulls the sampled region in from the lip edge — MediaPipe's inner lip landmarks sit slightly outside the real opening, and lips are bright. |
| blob grow/erode | Resizes the found blob. An open pass always runs first to despeckle. |
| tongue reject | Drops pixels red relative to their own brightness. Teeth are near-neutral; tongue is not. |
| prefer upper | Biases component choice toward the top of the cavity. Area alone picks the tongue when the mouth is wide. |
| teeth vertices | Radial sample count. |
| teeth avg ±f | Temporal average over the contour. |
| teeth dwell | Frames a presence change must persist before it takes effect. |
**Tongue** as its own part would work the same way, gated on redness instead of
brightness and biased low rather than high. Not implemented: it is not visible in
the test footage, which reads as a dark cavity with a bright upper-teeth band.
2026-09-24 15:12:09 -04:00
Registered photo underlay as the plate reference
The generated face oval was never going to be good enough to draw from:
MediaPipe's face oval is the FACE boundary, cut at the hairline and excluding
hair, ears, jaw and neck, so it is an egg by construction. Segmentation would
give a real head outline but costs a 16MB model and per-frame inference for a
shape that gets replaced by a drawing anyway.
So the plate layer becomes switchable, and the useful modes are photographic:
the source frame mapped into raster space through the same transform chain the
contours go through. Registration is the whole point - the head sits still and
a drawing traced from the underlay is already aligned to the mouth. An
unregistered underlay would be decoration.
- underlay.js: pixel->raster affine (a general affine, since MediaPipe
normalises x by width and y by height), registered draw, palette posterise
- plate modes: photo / photo dim / posterized / oval / oval+photo / none, B cycles
- worksheet cells are registered composites rather than raw crops
- Save frame 4x writes a 1280x800 PNG to draw on
- selftest: FACE_OVAL simplicity, which was never asserted; a wrong ordering
there reads as a lumpy plate rather than an obvious bowtie
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 14:57:59 -04:00
## The plate is reference, not art
The plate layer has several representations because its job changes. Cycle with
< kbd > B< / kbd > :
| Mode | For |
| --- | --- |
| photo dim / photo | **Drawing over.** The source frame, stabilised. |
| posterized | The footage quantised into the ramp — a look test asked of the source rather than of a drawing. |
| oval | A flat stand-in, to judge the mouth against something. |
| oval + photo | Checking the stand-in against the real head. |
| none | Mouth alone. |
Photo modes are **registered** : the frame is mapped into raster space through the
same transform chain the contours go through, so the head sits still and a
drawing traced from it is already aligned to the mouth. An unregistered underlay
would be decorative.
`MediaPipe` 's face oval is the *face* boundary — it cuts at the hairline and
excludes hair, ears, jaw underside and neck — so as a head silhouette it is an
egg by construction, and no landmark precision fixes that. Hence the photo.
**Save frame 4x** writes the current registered composite as a 1280× 800 PNG to
draw on.
2026-09-24 14:53:09 -04:00
## Two kinds of sparseness
Sparseness has two unrelated causes, and conflating them was the original design
error here. **Aesthetic** sparseness is set by the extraction rate — pick 12fps and
you have already chosen your timing. **Labour** sparseness is a human drawing
each one, and it binds only on the plate.
So the mouth keeps **every** frame: it is traced, and therefore free. In limited
animation lip sync is routinely the densest element, on 1s, while heads hold on
2s and 3s.
The frame strip is the editing surface for the other half: which frames need
their own plate drawing. Everything starts kept; delete what you don't want.
**Suggest** runs error-tolerance decimation over head pose as a starting point,
then you hand-correct.
2026-09-24 14:38:07 -04:00
2026-09-24 15:34:08 -04:00
**Mouth lead** is not a correction for a bug. A centred moving average has no
phase lag, so smoothing does not delay anything — but it blurs onsets, and the
visually salient moment of a mouth opening moves later even though the mean does
not. Animators also draw mouth shapes one or two frames ahead of the sound as
standard practice. Only the performance tracks shift; the head stays with the
audio, since it is the mouth that should anticipate. The lead is baked into the
exported take, so the renderer never needs to know about it.
2026-09-24 14:53:09 -04:00
`contour avg` is a deliberate, bounded exception to "never smooth the contour" in
`../docs/roto-puppet.md` . That rule held while keys were sparse, because sampling
at velocity minima rejected detector noise for free. With a key on every frame it
does not, so a radius shorter than the shortest articulation worth keeping is
justified — at 12fps, articulation spans 3– 6 frames and detector noise is
per-frame, so ±1 separates them and ±3 starts eating speech.
2026-09-24 14:38:07 -04:00
## Tests
```sh
chromium --headless --virtual-time-budget=8000 --dump-dom \
http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+'
```
2026-09-24 15:25:39 -04:00
Or open `selftest.html` . 41 assertions over the stages below detection, plus a
wiring cross-check: every `el('id')` in `app.js` must exist in `index.html` . A
knob wired in one but not the other throws during wiring, which aborts the rest
of the module and leaves a blank page — a symptom that points nowhere near its
cause, and which has happened twice.
2026-09-24 14:38:07 -04:00
The ring-simplicity check is the load-bearing one. Because `hold` parts *cut*
between poses instead of interpolating, a ring whose vertex order is wrong
self-intersects and renders as blocks meeting at corners — and it is invisible at
odd `verts/2` and obvious at even, so it needs an assertion rather than an
eyeball.
## Not done yet
2026-09-24 14:53:09 -04:00
Eyes and irises; hand-drawn head plates and per-plate mouth slots (the strip
decides *which frames need one* , but you cannot yet supply the drawing); real
2026-09-24 14:38:07 -04:00
performer→character calibration (currently identity, fitting the face oval to the
2026-09-24 14:53:09 -04:00
canvas); the override layer; anything on the Animator Pro side. The plate is a
face-oval polygon per kept frame — it exists so the mouth has a face to read
2026-09-24 14:38:07 -04:00
against, not to look good.