The generated face oval was never going to be good enough to draw from: MediaPipe's face oval is the FACE boundary, cut at the hairline and excluding hair, ears, jaw and neck, so it is an egg by construction. Segmentation would give a real head outline but costs a 16MB model and per-frame inference for a shape that gets replaced by a drawing anyway. So the plate layer becomes switchable, and the useful modes are photographic: the source frame mapped into raster space through the same transform chain the contours go through. Registration is the whole point - the head sits still and a drawing traced from the underlay is already aligned to the mouth. An unregistered underlay would be decoration. - underlay.js: pixel->raster affine (a general affine, since MediaPipe normalises x by width and y by height), registered draw, palette posterise - plate modes: photo / photo dim / posterized / oval / oval+photo / none, B cycles - worksheet cells are registered composites rather than raw crops - Save frame 4x writes a 1280x800 PNG to draw on - selftest: FACE_OVAL simplicity, which was never asserted; a wrong ordering there reads as a lumpy plate rather than an obvious bowtie Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
135 lines
5.8 KiB
Markdown
135 lines
5.8 KiB
Markdown
# roto
|
||
|
||
Video → **take file** builder for the Animator Pro rotoscope/puppet pipeline
|
||
described in `../docs/roto-puppet.md`.
|
||
|
||
This is the analysis and tuning half. It stabilises a face out of a clip, reduces
|
||
the lip contour to a handful of vertices, selects sparse keys on motion extremes,
|
||
and previews the result as flat indexed fills — so the *look and timing* can be
|
||
judged in seconds rather than through a minutes-long Animator Pro render. It
|
||
emits a `.take` file; nothing here touches Animator Pro.
|
||
|
||
All policy lives here. The take arrives at the renderer with its keys already
|
||
chosen.
|
||
|
||
## Run
|
||
|
||
```sh
|
||
python3 -m http.server 8777 # from this directory
|
||
# open http://127.0.0.1:8777
|
||
```
|
||
|
||
**Synthetic take** needs no video and exercises everything below detection.
|
||
|
||
For real footage:
|
||
|
||
```sh
|
||
./extract.sh /path/to/clip.mov 12 # -> frames/*.png, audio.wav, manifest.json
|
||
```
|
||
|
||
then **Load frames**. MediaPipe's wasm is fetched from jsdelivr on first use;
|
||
`face_landmarker.task` is local.
|
||
|
||
Frames are pre-extracted rather than decoded in the page because browser video
|
||
seeking is approximate and `requestVideoFrameCallback` only delivers frames at
|
||
playback speed — neither gives a deterministic per-frame pass.
|
||
|
||
`manifest.json` records the true extraction rate. The page reads it rather than
|
||
assuming, because a guessed fps desynchronises audio from picture — and sync is
|
||
the one thing this view exists to show.
|
||
|
||
**Audio is the playback clock**: `frame = floor(audio.currentTime * fps)`. A slow
|
||
render loop therefore drops frames instead of drifting, and ½x / ¼x work by
|
||
setting `playbackRate` with the picture following for free.
|
||
|
||
## Shooting for it
|
||
|
||
Near-frontal, good light, consistent scale, head reasonably still. Hold a neutral
|
||
closed mouth for a second at the top of the take: that frame is picked
|
||
automatically as the neutral and drives calibration and the placeholder plate.
|
||
|
||
Out-of-plane head rotation cannot be stabilised away by a 2D similarity
|
||
transform — the `residual` readout rises when it happens. The pipeline's answer
|
||
is hand-drawn head plates, which this tool does not yet do.
|
||
|
||
## Knobs
|
||
|
||
| Knob | What it does |
|
||
| --- | --- |
|
||
| vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. |
|
||
| contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. |
|
||
| anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform*. |
|
||
| closed-mouth cut | Aperture below which the mouth interior is emitted as `hidden`. |
|
||
| suggest tolerance | Max head movement before a new plate drawing is required. Affects **Suggest** only. |
|
||
|
||
## The plate is reference, not art
|
||
|
||
The plate layer has several representations because its job changes. Cycle with
|
||
<kbd>B</kbd>:
|
||
|
||
| Mode | For |
|
||
| --- | --- |
|
||
| photo dim / photo | **Drawing over.** The source frame, stabilised. |
|
||
| posterized | The footage quantised into the ramp — a look test asked of the source rather than of a drawing. |
|
||
| oval | A flat stand-in, to judge the mouth against something. |
|
||
| oval + photo | Checking the stand-in against the real head. |
|
||
| none | Mouth alone. |
|
||
|
||
Photo modes are **registered**: the frame is mapped into raster space through the
|
||
same transform chain the contours go through, so the head sits still and a
|
||
drawing traced from it is already aligned to the mouth. An unregistered underlay
|
||
would be decorative.
|
||
|
||
`MediaPipe`'s face oval is the *face* boundary — it cuts at the hairline and
|
||
excludes hair, ears, jaw underside and neck — so as a head silhouette it is an
|
||
egg by construction, and no landmark precision fixes that. Hence the photo.
|
||
|
||
**Save frame 4x** writes the current registered composite as a 1280×800 PNG to
|
||
draw on.
|
||
|
||
## Two kinds of sparseness
|
||
|
||
Sparseness has two unrelated causes, and conflating them was the original design
|
||
error here. **Aesthetic** sparseness is set by the extraction rate — pick 12fps and
|
||
you have already chosen your timing. **Labour** sparseness is a human drawing
|
||
each one, and it binds only on the plate.
|
||
|
||
So the mouth keeps **every** frame: it is traced, and therefore free. In limited
|
||
animation lip sync is routinely the densest element, on 1s, while heads hold on
|
||
2s and 3s.
|
||
|
||
The frame strip is the editing surface for the other half: which frames need
|
||
their own plate drawing. Everything starts kept; delete what you don't want.
|
||
**Suggest** runs error-tolerance decimation over head pose as a starting point,
|
||
then you hand-correct.
|
||
|
||
`contour avg` is a deliberate, bounded exception to "never smooth the contour" in
|
||
`../docs/roto-puppet.md`. That rule held while keys were sparse, because sampling
|
||
at velocity minima rejected detector noise for free. With a key on every frame it
|
||
does not, so a radius shorter than the shortest articulation worth keeping is
|
||
justified — at 12fps, articulation spans 3–6 frames and detector noise is
|
||
per-frame, so ±1 separates them and ±3 starts eating speech.
|
||
|
||
## Tests
|
||
|
||
```sh
|
||
chromium --headless --virtual-time-budget=8000 --dump-dom \
|
||
http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+'
|
||
```
|
||
|
||
Or open `selftest.html`. 29 assertions over the stages below detection.
|
||
|
||
The ring-simplicity check is the load-bearing one. Because `hold` parts *cut*
|
||
between poses instead of interpolating, a ring whose vertex order is wrong
|
||
self-intersects and renders as blocks meeting at corners — and it is invisible at
|
||
odd `verts/2` and obvious at even, so it needs an assertion rather than an
|
||
eyeball.
|
||
|
||
## Not done yet
|
||
|
||
Eyes and irises; hand-drawn head plates and per-plate mouth slots (the strip
|
||
decides *which frames need one*, but you cannot yet supply the drawing); real
|
||
performer→character calibration (currently identity, fitting the face oval to the
|
||
canvas); the override layer; anything on the Animator Pro side. The plate is a
|
||
face-oval polygon per kept frame — it exists so the mouth has a face to read
|
||
against, not to look good.
|