2026-09-24 14:38:07 -04:00
|
|
|
|
# roto
|
|
|
|
|
|
|
|
|
|
|
|
Video → **take file** builder for the Animator Pro rotoscope/puppet pipeline
|
|
|
|
|
|
described in `../docs/roto-puppet.md`.
|
|
|
|
|
|
|
|
|
|
|
|
This is the analysis and tuning half. It stabilises a face out of a clip, reduces
|
|
|
|
|
|
the lip contour to a handful of vertices, selects sparse keys on motion extremes,
|
|
|
|
|
|
and previews the result as flat indexed fills — so the *look and timing* can be
|
|
|
|
|
|
judged in seconds rather than through a minutes-long Animator Pro render. It
|
|
|
|
|
|
emits a `.take` file; nothing here touches Animator Pro.
|
|
|
|
|
|
|
|
|
|
|
|
All policy lives here. The take arrives at the renderer with its keys already
|
|
|
|
|
|
chosen.
|
|
|
|
|
|
|
|
|
|
|
|
## Run
|
|
|
|
|
|
|
|
|
|
|
|
```sh
|
|
|
|
|
|
python3 -m http.server 8777 # from this directory
|
|
|
|
|
|
# open http://127.0.0.1:8777
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
**Synthetic take** needs no video and exercises everything below detection.
|
|
|
|
|
|
|
|
|
|
|
|
For real footage:
|
|
|
|
|
|
|
|
|
|
|
|
```sh
|
2026-09-24 14:53:09 -04:00
|
|
|
|
./extract.sh /path/to/clip.mov 12 # -> frames/*.png, audio.wav, manifest.json
|
2026-09-24 14:38:07 -04:00
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
then **Load frames**. MediaPipe's wasm is fetched from jsdelivr on first use;
|
|
|
|
|
|
`face_landmarker.task` is local.
|
|
|
|
|
|
|
|
|
|
|
|
Frames are pre-extracted rather than decoded in the page because browser video
|
|
|
|
|
|
seeking is approximate and `requestVideoFrameCallback` only delivers frames at
|
|
|
|
|
|
playback speed — neither gives a deterministic per-frame pass.
|
|
|
|
|
|
|
2026-09-24 14:53:09 -04:00
|
|
|
|
`manifest.json` records the true extraction rate. The page reads it rather than
|
|
|
|
|
|
assuming, because a guessed fps desynchronises audio from picture — and sync is
|
|
|
|
|
|
the one thing this view exists to show.
|
|
|
|
|
|
|
|
|
|
|
|
**Audio is the playback clock**: `frame = floor(audio.currentTime * fps)`. A slow
|
|
|
|
|
|
render loop therefore drops frames instead of drifting, and ½x / ¼x work by
|
|
|
|
|
|
setting `playbackRate` with the picture following for free.
|
|
|
|
|
|
|
2026-09-24 14:38:07 -04:00
|
|
|
|
## Shooting for it
|
|
|
|
|
|
|
|
|
|
|
|
Near-frontal, good light, consistent scale, head reasonably still. Hold a neutral
|
|
|
|
|
|
closed mouth for a second at the top of the take: that frame is picked
|
|
|
|
|
|
automatically as the neutral and drives calibration and the placeholder plate.
|
|
|
|
|
|
|
|
|
|
|
|
Out-of-plane head rotation cannot be stabilised away by a 2D similarity
|
|
|
|
|
|
transform — the `residual` readout rises when it happens. The pipeline's answer
|
|
|
|
|
|
is hand-drawn head plates, which this tool does not yet do.
|
|
|
|
|
|
|
|
|
|
|
|
## Knobs
|
|
|
|
|
|
|
|
|
|
|
|
| Knob | What it does |
|
|
|
|
|
|
| --- | --- |
|
|
|
|
|
|
| vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. |
|
2026-09-24 14:53:09 -04:00
|
|
|
|
| contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. |
|
|
|
|
|
|
| anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform*. |
|
2026-09-24 14:38:07 -04:00
|
|
|
|
| closed-mouth cut | Aperture below which the mouth interior is emitted as `hidden`. |
|
2026-09-24 14:53:09 -04:00
|
|
|
|
| suggest tolerance | Max head movement before a new plate drawing is required. Affects **Suggest** only. |
|
|
|
|
|
|
|
2026-09-24 15:12:09 -04:00
|
|
|
|
## Teeth
|
|
|
|
|
|
|
|
|
|
|
|
MediaPipe has no landmarks inside the lips: the inner ring bounds the cavity and
|
|
|
|
|
|
everything within it is just pixels. So teeth come from the image — but tracing
|
|
|
|
|
|
the bright blob would produce a new contour every frame with no vertex
|
|
|
|
|
|
correspondence, which is precisely the boil the design exists to avoid.
|
|
|
|
|
|
|
|
|
|
|
|
So the extraction yields a **scalar, not a shape**. The teeth polygon is the inner
|
|
|
|
|
|
lip ring clipped to a horizontal line, and only that line's height is measured
|
|
|
|
|
|
(Otsu threshold within the cavity, scanned from the top). The silhouette is
|
|
|
|
|
|
therefore always the mouth's own shape — stable by construction — and the only
|
|
|
|
|
|
per-frame variable is one number, which smooths trivially. It is also how the
|
|
|
|
|
|
shape gets drawn by hand: a band bounded by the lip.
|
|
|
|
|
|
|
|
|
|
|
|
Presence uses hysteresis plus a minimum dwell, the same treatment plate selection
|
|
|
|
|
|
gets, because a teeth block that blinks on and off for single frames is worse
|
|
|
|
|
|
than one that is simply absent.
|
|
|
|
|
|
|
|
|
|
|
|
**Tongue** would work the same way — a shape filling the lower cavity, gated on a
|
|
|
|
|
|
redness rather than a brightness statistic. Not implemented, because it is not
|
|
|
|
|
|
visible in the test footage: the cavity reads as dark with a bright upper-teeth
|
|
|
|
|
|
band and nothing else.
|
|
|
|
|
|
|
Registered photo underlay as the plate reference
The generated face oval was never going to be good enough to draw from:
MediaPipe's face oval is the FACE boundary, cut at the hairline and excluding
hair, ears, jaw and neck, so it is an egg by construction. Segmentation would
give a real head outline but costs a 16MB model and per-frame inference for a
shape that gets replaced by a drawing anyway.
So the plate layer becomes switchable, and the useful modes are photographic:
the source frame mapped into raster space through the same transform chain the
contours go through. Registration is the whole point - the head sits still and
a drawing traced from the underlay is already aligned to the mouth. An
unregistered underlay would be decoration.
- underlay.js: pixel->raster affine (a general affine, since MediaPipe
normalises x by width and y by height), registered draw, palette posterise
- plate modes: photo / photo dim / posterized / oval / oval+photo / none, B cycles
- worksheet cells are registered composites rather than raw crops
- Save frame 4x writes a 1280x800 PNG to draw on
- selftest: FACE_OVAL simplicity, which was never asserted; a wrong ordering
there reads as a lumpy plate rather than an obvious bowtie
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 14:57:59 -04:00
|
|
|
|
## The plate is reference, not art
|
|
|
|
|
|
|
|
|
|
|
|
The plate layer has several representations because its job changes. Cycle with
|
|
|
|
|
|
<kbd>B</kbd>:
|
|
|
|
|
|
|
|
|
|
|
|
| Mode | For |
|
|
|
|
|
|
| --- | --- |
|
|
|
|
|
|
| photo dim / photo | **Drawing over.** The source frame, stabilised. |
|
|
|
|
|
|
| posterized | The footage quantised into the ramp — a look test asked of the source rather than of a drawing. |
|
|
|
|
|
|
| oval | A flat stand-in, to judge the mouth against something. |
|
|
|
|
|
|
| oval + photo | Checking the stand-in against the real head. |
|
|
|
|
|
|
| none | Mouth alone. |
|
|
|
|
|
|
|
|
|
|
|
|
Photo modes are **registered**: the frame is mapped into raster space through the
|
|
|
|
|
|
same transform chain the contours go through, so the head sits still and a
|
|
|
|
|
|
drawing traced from it is already aligned to the mouth. An unregistered underlay
|
|
|
|
|
|
would be decorative.
|
|
|
|
|
|
|
|
|
|
|
|
`MediaPipe`'s face oval is the *face* boundary — it cuts at the hairline and
|
|
|
|
|
|
excludes hair, ears, jaw underside and neck — so as a head silhouette it is an
|
|
|
|
|
|
egg by construction, and no landmark precision fixes that. Hence the photo.
|
|
|
|
|
|
|
|
|
|
|
|
**Save frame 4x** writes the current registered composite as a 1280×800 PNG to
|
|
|
|
|
|
draw on.
|
|
|
|
|
|
|
2026-09-24 14:53:09 -04:00
|
|
|
|
## Two kinds of sparseness
|
|
|
|
|
|
|
|
|
|
|
|
Sparseness has two unrelated causes, and conflating them was the original design
|
|
|
|
|
|
error here. **Aesthetic** sparseness is set by the extraction rate — pick 12fps and
|
|
|
|
|
|
you have already chosen your timing. **Labour** sparseness is a human drawing
|
|
|
|
|
|
each one, and it binds only on the plate.
|
|
|
|
|
|
|
|
|
|
|
|
So the mouth keeps **every** frame: it is traced, and therefore free. In limited
|
|
|
|
|
|
animation lip sync is routinely the densest element, on 1s, while heads hold on
|
|
|
|
|
|
2s and 3s.
|
|
|
|
|
|
|
|
|
|
|
|
The frame strip is the editing surface for the other half: which frames need
|
|
|
|
|
|
their own plate drawing. Everything starts kept; delete what you don't want.
|
|
|
|
|
|
**Suggest** runs error-tolerance decimation over head pose as a starting point,
|
|
|
|
|
|
then you hand-correct.
|
2026-09-24 14:38:07 -04:00
|
|
|
|
|
2026-09-24 14:53:09 -04:00
|
|
|
|
`contour avg` is a deliberate, bounded exception to "never smooth the contour" in
|
|
|
|
|
|
`../docs/roto-puppet.md`. That rule held while keys were sparse, because sampling
|
|
|
|
|
|
at velocity minima rejected detector noise for free. With a key on every frame it
|
|
|
|
|
|
does not, so a radius shorter than the shortest articulation worth keeping is
|
|
|
|
|
|
justified — at 12fps, articulation spans 3–6 frames and detector noise is
|
|
|
|
|
|
per-frame, so ±1 separates them and ±3 starts eating speech.
|
2026-09-24 14:38:07 -04:00
|
|
|
|
|
|
|
|
|
|
## Tests
|
|
|
|
|
|
|
|
|
|
|
|
```sh
|
|
|
|
|
|
chromium --headless --virtual-time-budget=8000 --dump-dom \
|
|
|
|
|
|
http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+'
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
Or open `selftest.html`. 29 assertions over the stages below detection.
|
|
|
|
|
|
|
|
|
|
|
|
The ring-simplicity check is the load-bearing one. Because `hold` parts *cut*
|
|
|
|
|
|
between poses instead of interpolating, a ring whose vertex order is wrong
|
|
|
|
|
|
self-intersects and renders as blocks meeting at corners — and it is invisible at
|
|
|
|
|
|
odd `verts/2` and obvious at even, so it needs an assertion rather than an
|
|
|
|
|
|
eyeball.
|
|
|
|
|
|
|
|
|
|
|
|
## Not done yet
|
|
|
|
|
|
|
2026-09-24 14:53:09 -04:00
|
|
|
|
Eyes and irises; hand-drawn head plates and per-plate mouth slots (the strip
|
|
|
|
|
|
decides *which frames need one*, but you cannot yet supply the drawing); real
|
2026-09-24 14:38:07 -04:00
|
|
|
|
performer→character calibration (currently identity, fitting the face oval to the
|
2026-09-24 14:53:09 -04:00
|
|
|
|
canvas); the override layer; anything on the Animator Pro side. The plate is a
|
|
|
|
|
|
face-oval polygon per kept frame — it exists so the mouth has a face to read
|
2026-09-24 14:38:07 -04:00
|
|
|
|
against, not to look good.
|