MediaPipe has no landmarks inside the lips, so teeth have to come from pixels. Tracing the bright blob would give a new contour every frame with no vertex correspondence - the exact boil docs/roto-puppet.md warns about. So the measurement yields a scalar, not a shape: Otsu within the cavity, scanned from the top for where the bright run stops, giving one line height per frame. The teeth polygon is the inner lip ring clipped to that line, so the silhouette is always the mouth's own shape and cannot disagree with the lips around it, and the only per-frame variable is a single number that smooths trivially. Presence gets hysteresis and minimum dwell, as plate selection does: a teeth block blinking on and off for single frames is worse than one simply absent. Tongue is not implemented. The same scalar approach would apply, gated on redness rather than brightness, but it is not visible in the test footage - the cavity reads dark with a bright upper-teeth band and nothing else. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
158 lines
7 KiB
Markdown
158 lines
7 KiB
Markdown
# roto
|
||
|
||
Video → **take file** builder for the Animator Pro rotoscope/puppet pipeline
|
||
described in `../docs/roto-puppet.md`.
|
||
|
||
This is the analysis and tuning half. It stabilises a face out of a clip, reduces
|
||
the lip contour to a handful of vertices, selects sparse keys on motion extremes,
|
||
and previews the result as flat indexed fills — so the *look and timing* can be
|
||
judged in seconds rather than through a minutes-long Animator Pro render. It
|
||
emits a `.take` file; nothing here touches Animator Pro.
|
||
|
||
All policy lives here. The take arrives at the renderer with its keys already
|
||
chosen.
|
||
|
||
## Run
|
||
|
||
```sh
|
||
python3 -m http.server 8777 # from this directory
|
||
# open http://127.0.0.1:8777
|
||
```
|
||
|
||
**Synthetic take** needs no video and exercises everything below detection.
|
||
|
||
For real footage:
|
||
|
||
```sh
|
||
./extract.sh /path/to/clip.mov 12 # -> frames/*.png, audio.wav, manifest.json
|
||
```
|
||
|
||
then **Load frames**. MediaPipe's wasm is fetched from jsdelivr on first use;
|
||
`face_landmarker.task` is local.
|
||
|
||
Frames are pre-extracted rather than decoded in the page because browser video
|
||
seeking is approximate and `requestVideoFrameCallback` only delivers frames at
|
||
playback speed — neither gives a deterministic per-frame pass.
|
||
|
||
`manifest.json` records the true extraction rate. The page reads it rather than
|
||
assuming, because a guessed fps desynchronises audio from picture — and sync is
|
||
the one thing this view exists to show.
|
||
|
||
**Audio is the playback clock**: `frame = floor(audio.currentTime * fps)`. A slow
|
||
render loop therefore drops frames instead of drifting, and ½x / ¼x work by
|
||
setting `playbackRate` with the picture following for free.
|
||
|
||
## Shooting for it
|
||
|
||
Near-frontal, good light, consistent scale, head reasonably still. Hold a neutral
|
||
closed mouth for a second at the top of the take: that frame is picked
|
||
automatically as the neutral and drives calibration and the placeholder plate.
|
||
|
||
Out-of-plane head rotation cannot be stabilised away by a 2D similarity
|
||
transform — the `residual` readout rises when it happens. The pipeline's answer
|
||
is hand-drawn head plates, which this tool does not yet do.
|
||
|
||
## Knobs
|
||
|
||
| Knob | What it does |
|
||
| --- | --- |
|
||
| vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. |
|
||
| contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. |
|
||
| anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform*. |
|
||
| closed-mouth cut | Aperture below which the mouth interior is emitted as `hidden`. |
|
||
| suggest tolerance | Max head movement before a new plate drawing is required. Affects **Suggest** only. |
|
||
|
||
## Teeth
|
||
|
||
MediaPipe has no landmarks inside the lips: the inner ring bounds the cavity and
|
||
everything within it is just pixels. So teeth come from the image — but tracing
|
||
the bright blob would produce a new contour every frame with no vertex
|
||
correspondence, which is precisely the boil the design exists to avoid.
|
||
|
||
So the extraction yields a **scalar, not a shape**. The teeth polygon is the inner
|
||
lip ring clipped to a horizontal line, and only that line's height is measured
|
||
(Otsu threshold within the cavity, scanned from the top). The silhouette is
|
||
therefore always the mouth's own shape — stable by construction — and the only
|
||
per-frame variable is one number, which smooths trivially. It is also how the
|
||
shape gets drawn by hand: a band bounded by the lip.
|
||
|
||
Presence uses hysteresis plus a minimum dwell, the same treatment plate selection
|
||
gets, because a teeth block that blinks on and off for single frames is worse
|
||
than one that is simply absent.
|
||
|
||
**Tongue** would work the same way — a shape filling the lower cavity, gated on a
|
||
redness rather than a brightness statistic. Not implemented, because it is not
|
||
visible in the test footage: the cavity reads as dark with a bright upper-teeth
|
||
band and nothing else.
|
||
|
||
## The plate is reference, not art
|
||
|
||
The plate layer has several representations because its job changes. Cycle with
|
||
<kbd>B</kbd>:
|
||
|
||
| Mode | For |
|
||
| --- | --- |
|
||
| photo dim / photo | **Drawing over.** The source frame, stabilised. |
|
||
| posterized | The footage quantised into the ramp — a look test asked of the source rather than of a drawing. |
|
||
| oval | A flat stand-in, to judge the mouth against something. |
|
||
| oval + photo | Checking the stand-in against the real head. |
|
||
| none | Mouth alone. |
|
||
|
||
Photo modes are **registered**: the frame is mapped into raster space through the
|
||
same transform chain the contours go through, so the head sits still and a
|
||
drawing traced from it is already aligned to the mouth. An unregistered underlay
|
||
would be decorative.
|
||
|
||
`MediaPipe`'s face oval is the *face* boundary — it cuts at the hairline and
|
||
excludes hair, ears, jaw underside and neck — so as a head silhouette it is an
|
||
egg by construction, and no landmark precision fixes that. Hence the photo.
|
||
|
||
**Save frame 4x** writes the current registered composite as a 1280×800 PNG to
|
||
draw on.
|
||
|
||
## Two kinds of sparseness
|
||
|
||
Sparseness has two unrelated causes, and conflating them was the original design
|
||
error here. **Aesthetic** sparseness is set by the extraction rate — pick 12fps and
|
||
you have already chosen your timing. **Labour** sparseness is a human drawing
|
||
each one, and it binds only on the plate.
|
||
|
||
So the mouth keeps **every** frame: it is traced, and therefore free. In limited
|
||
animation lip sync is routinely the densest element, on 1s, while heads hold on
|
||
2s and 3s.
|
||
|
||
The frame strip is the editing surface for the other half: which frames need
|
||
their own plate drawing. Everything starts kept; delete what you don't want.
|
||
**Suggest** runs error-tolerance decimation over head pose as a starting point,
|
||
then you hand-correct.
|
||
|
||
`contour avg` is a deliberate, bounded exception to "never smooth the contour" in
|
||
`../docs/roto-puppet.md`. That rule held while keys were sparse, because sampling
|
||
at velocity minima rejected detector noise for free. With a key on every frame it
|
||
does not, so a radius shorter than the shortest articulation worth keeping is
|
||
justified — at 12fps, articulation spans 3–6 frames and detector noise is
|
||
per-frame, so ±1 separates them and ±3 starts eating speech.
|
||
|
||
## Tests
|
||
|
||
```sh
|
||
chromium --headless --virtual-time-budget=8000 --dump-dom \
|
||
http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+'
|
||
```
|
||
|
||
Or open `selftest.html`. 29 assertions over the stages below detection.
|
||
|
||
The ring-simplicity check is the load-bearing one. Because `hold` parts *cut*
|
||
between poses instead of interpolating, a ring whose vertex order is wrong
|
||
self-intersects and renders as blocks meeting at corners — and it is invisible at
|
||
odd `verts/2` and obvious at even, so it needs an assertion rather than an
|
||
eyeball.
|
||
|
||
## Not done yet
|
||
|
||
Eyes and irises; hand-drawn head plates and per-plate mouth slots (the strip
|
||
decides *which frames need one*, but you cannot yet supply the drawing); real
|
||
performer→character calibration (currently identity, fitting the face oval to the
|
||
canvas); the override layer; anything on the Animator Pro side. The plate is a
|
||
face-oval polygon per kept frame — it exists so the mouth has a face to read
|
||
against, not to look good.
|