A knob wired in app.js but missing from the markup throws during wiring, aborts the module and leaves a blank page - a symptom pointing nowhere near its cause. It has now happened twice, so it gets a check rather than vigilance: the selftest fetches both files and compares the id sets in each direction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
170 lines
8 KiB
Markdown
170 lines
8 KiB
Markdown
# roto
|
||
|
||
Video → **take file** builder for the Animator Pro rotoscope/puppet pipeline
|
||
described in `../docs/roto-puppet.md`.
|
||
|
||
This is the analysis and tuning half. It stabilises a face out of a clip, reduces
|
||
the lip contour to a handful of vertices, selects sparse keys on motion extremes,
|
||
and previews the result as flat indexed fills — so the *look and timing* can be
|
||
judged in seconds rather than through a minutes-long Animator Pro render. It
|
||
emits a `.take` file; nothing here touches Animator Pro.
|
||
|
||
All policy lives here. The take arrives at the renderer with its keys already
|
||
chosen.
|
||
|
||
## Run
|
||
|
||
```sh
|
||
python3 -m http.server 8777 # from this directory
|
||
# open http://127.0.0.1:8777
|
||
```
|
||
|
||
**Synthetic take** needs no video and exercises everything below detection.
|
||
|
||
For real footage:
|
||
|
||
```sh
|
||
./extract.sh /path/to/clip.mov 12 # -> frames/*.png, audio.wav, manifest.json
|
||
```
|
||
|
||
then **Load frames**. MediaPipe's wasm is fetched from jsdelivr on first use;
|
||
`face_landmarker.task` is local.
|
||
|
||
Frames are pre-extracted rather than decoded in the page because browser video
|
||
seeking is approximate and `requestVideoFrameCallback` only delivers frames at
|
||
playback speed — neither gives a deterministic per-frame pass.
|
||
|
||
`manifest.json` records the true extraction rate. The page reads it rather than
|
||
assuming, because a guessed fps desynchronises audio from picture — and sync is
|
||
the one thing this view exists to show.
|
||
|
||
**Audio is the playback clock**: `frame = floor(audio.currentTime * fps)`. A slow
|
||
render loop therefore drops frames instead of drifting, and ½x / ¼x work by
|
||
setting `playbackRate` with the picture following for free.
|
||
|
||
## Shooting for it
|
||
|
||
Near-frontal, good light, consistent scale, head reasonably still. Hold a neutral
|
||
closed mouth for a second at the top of the take: that frame is picked
|
||
automatically as the neutral and drives calibration and the placeholder plate.
|
||
|
||
Out-of-plane head rotation cannot be stabilised away by a 2D similarity
|
||
transform — the `residual` readout rises when it happens. The pipeline's answer
|
||
is hand-drawn head plates, which this tool does not yet do.
|
||
|
||
## Knobs
|
||
|
||
| Knob | What it does |
|
||
| --- | --- |
|
||
| vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. |
|
||
| contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. |
|
||
| anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform*. |
|
||
| closed-mouth cut | Aperture below which the mouth interior is emitted as `hidden`. |
|
||
| suggest tolerance | Max head movement before a new plate drawing is required. Affects **Suggest** only. |
|
||
|
||
## Teeth
|
||
|
||
MediaPipe has no landmarks inside the lips: the inner ring bounds the cavity and
|
||
everything within it is just pixels. So teeth come from the image.
|
||
|
||
The hazard is vertex correspondence — a traced contour reorders between frames
|
||
and boils. The way out for a blob specifically is **radial sampling**: march
|
||
outward from the blob's centroid along N fixed directions and take the last pixel
|
||
inside. Vertex *k* is then always "the extent in direction *k*", so correspondence
|
||
holds by construction, the vertex count is fixed, and temporal smoothing is well
|
||
defined with no reordering possible. It also produces a star-shaped reduction,
|
||
which is what flat blocks of colour want.
|
||
|
||
The **teeth measurement** panel shows exactly what is sampled: region dimmed,
|
||
kept pixels green, extracted contour amber. Tune against that, not the numbers.
|
||
|
||
| Knob | What it does |
|
||
| --- | --- |
|
||
| teeth contrast | Gate on the separation between the cavity's dark and bright class means. Otsu always returns *some* threshold, so this is what stops it inventing teeth in a dark mouth. |
|
||
| cavity erode | Pulls the sampled region in from the lip edge — MediaPipe's inner lip landmarks sit slightly outside the real opening, and lips are bright. |
|
||
| blob grow/erode | Resizes the found blob. An open pass always runs first to despeckle. |
|
||
| tongue reject | Drops pixels red relative to their own brightness. Teeth are near-neutral; tongue is not. |
|
||
| prefer upper | Biases component choice toward the top of the cavity. Area alone picks the tongue when the mouth is wide. |
|
||
| teeth vertices | Radial sample count. |
|
||
| teeth avg ±f | Temporal average over the contour. |
|
||
| teeth dwell | Frames a presence change must persist before it takes effect. |
|
||
|
||
**Tongue** as its own part would work the same way, gated on redness instead of
|
||
brightness and biased low rather than high. Not implemented: it is not visible in
|
||
the test footage, which reads as a dark cavity with a bright upper-teeth band.
|
||
|
||
## The plate is reference, not art
|
||
|
||
The plate layer has several representations because its job changes. Cycle with
|
||
<kbd>B</kbd>:
|
||
|
||
| Mode | For |
|
||
| --- | --- |
|
||
| photo dim / photo | **Drawing over.** The source frame, stabilised. |
|
||
| posterized | The footage quantised into the ramp — a look test asked of the source rather than of a drawing. |
|
||
| oval | A flat stand-in, to judge the mouth against something. |
|
||
| oval + photo | Checking the stand-in against the real head. |
|
||
| none | Mouth alone. |
|
||
|
||
Photo modes are **registered**: the frame is mapped into raster space through the
|
||
same transform chain the contours go through, so the head sits still and a
|
||
drawing traced from it is already aligned to the mouth. An unregistered underlay
|
||
would be decorative.
|
||
|
||
`MediaPipe`'s face oval is the *face* boundary — it cuts at the hairline and
|
||
excludes hair, ears, jaw underside and neck — so as a head silhouette it is an
|
||
egg by construction, and no landmark precision fixes that. Hence the photo.
|
||
|
||
**Save frame 4x** writes the current registered composite as a 1280×800 PNG to
|
||
draw on.
|
||
|
||
## Two kinds of sparseness
|
||
|
||
Sparseness has two unrelated causes, and conflating them was the original design
|
||
error here. **Aesthetic** sparseness is set by the extraction rate — pick 12fps and
|
||
you have already chosen your timing. **Labour** sparseness is a human drawing
|
||
each one, and it binds only on the plate.
|
||
|
||
So the mouth keeps **every** frame: it is traced, and therefore free. In limited
|
||
animation lip sync is routinely the densest element, on 1s, while heads hold on
|
||
2s and 3s.
|
||
|
||
The frame strip is the editing surface for the other half: which frames need
|
||
their own plate drawing. Everything starts kept; delete what you don't want.
|
||
**Suggest** runs error-tolerance decimation over head pose as a starting point,
|
||
then you hand-correct.
|
||
|
||
`contour avg` is a deliberate, bounded exception to "never smooth the contour" in
|
||
`../docs/roto-puppet.md`. That rule held while keys were sparse, because sampling
|
||
at velocity minima rejected detector noise for free. With a key on every frame it
|
||
does not, so a radius shorter than the shortest articulation worth keeping is
|
||
justified — at 12fps, articulation spans 3–6 frames and detector noise is
|
||
per-frame, so ±1 separates them and ±3 starts eating speech.
|
||
|
||
## Tests
|
||
|
||
```sh
|
||
chromium --headless --virtual-time-budget=8000 --dump-dom \
|
||
http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+'
|
||
```
|
||
|
||
Or open `selftest.html`. 41 assertions over the stages below detection, plus a
|
||
wiring cross-check: every `el('id')` in `app.js` must exist in `index.html`. A
|
||
knob wired in one but not the other throws during wiring, which aborts the rest
|
||
of the module and leaves a blank page — a symptom that points nowhere near its
|
||
cause, and which has happened twice.
|
||
|
||
The ring-simplicity check is the load-bearing one. Because `hold` parts *cut*
|
||
between poses instead of interpolating, a ring whose vertex order is wrong
|
||
self-intersects and renders as blocks meeting at corners — and it is invisible at
|
||
odd `verts/2` and obvious at even, so it needs an assertion rather than an
|
||
eyeball.
|
||
|
||
## Not done yet
|
||
|
||
Eyes and irises; hand-drawn head plates and per-plate mouth slots (the strip
|
||
decides *which frames need one*, but you cannot yet supply the drawing); real
|
||
performer→character calibration (currently identity, fitting the face oval to the
|
||
canvas); the override layer; anything on the Animator Pro side. The plate is a
|
||
face-oval polygon per kept frame — it exists so the mouth has a face to read
|
||
against, not to look good.
|