Analysis half of the pipeline in docs/roto-puppet.md. Stabilises a face out of a clip via a similarity fit on rigid landmarks, reduces the lip contour to a fixed vertex budget, selects sparse keys on velocity minima, and previews the result as flat indexed fills so timing can be judged without an Animator Pro render. - landmarks.js ordered lip/oval rings; slot position is vertex identity - mathutil.js closed-form 2D similarity, Procrustes mean, transform smoothing - pipeline.js stabilise -> subsample -> key-select - raster.js indexed scanline fill, no antialiasing - take.js take-file writer - selftest.js 29 assertions, incl. ring simplicity at every vertex budget Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|---|---|---|
| js | ||
| .gitignore | ||
| extract.sh | ||
| index.html | ||
| README.md | ||
| selftest.html | ||
roto
Video → take file builder for the Animator Pro rotoscope/puppet pipeline
described in ../docs/roto-puppet.md.
This is the analysis and tuning half. It stabilises a face out of a clip, reduces
the lip contour to a handful of vertices, selects sparse keys on motion extremes,
and previews the result as flat indexed fills — so the look and timing can be
judged in seconds rather than through a minutes-long Animator Pro render. It
emits a .take file; nothing here touches Animator Pro.
All policy lives here. The take arrives at the renderer with its keys already chosen.
Run
python3 -m http.server 8777 # from this directory
# open http://127.0.0.1:8777
Synthetic take needs no video and exercises everything below detection.
For real footage:
./extract.sh /path/to/clip.mp4 24 # -> frames/0001.png …
then Load frames. MediaPipe's wasm is fetched from jsdelivr on first use;
face_landmarker.task is local.
Frames are pre-extracted rather than decoded in the page because browser video
seeking is approximate and requestVideoFrameCallback only delivers frames at
playback speed — neither gives a deterministic per-frame pass.
Shooting for it
Near-frontal, good light, consistent scale, head reasonably still. Hold a neutral closed mouth for a second at the top of the take: that frame is picked automatically as the neutral and drives calibration and the placeholder plate.
Out-of-plane head rotation cannot be stabilised away by a 2D similarity
transform — the residual readout rises when it happens. The pipeline's answer
is hand-drawn head plates, which this tool does not yet do.
Knobs
| Knob | What it does |
|---|---|
| vertices | Lip vertex budget. The reduction past what the footage supports is the style. |
| min hold | Minimum frames between keys. |
| change gate | Mean vertex movement required before a new key is accepted. |
| vel smoothing | Window on the velocity signal used to find extremes. |
| anchor smoothing | Window on the four similarity parameters. Smooths the transform, never the contour. |
| exposure | Grid that key frames snap onto. |
| closed-mouth cut | Aperture below which the mouth interior is emitted as hidden. |
Keys go on velocity minima, not distance thresholds: a threshold fires at the
frame it was crossed — partway through a transition — so poses land mushy and
late. The timeline shows the velocity curve, candidate minima, accepted keys
(f) and their pre-snap extremes (src); a large f/src gap means min-hold
and exposure are fighting.
Tests
chromium --headless --virtual-time-budget=8000 --dump-dom \
http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+'
Or open selftest.html. 29 assertions over the stages below detection.
The ring-simplicity check is the load-bearing one. Because hold parts cut
between poses instead of interpolating, a ring whose vertex order is wrong
self-intersects and renders as blocks meeting at corners — and it is invisible at
odd verts/2 and obvious at even, so it needs an assertion rather than an
eyeball.
Not done yet
Eyes and irises; hand-drawn head plates and per-plate mouth slots; real performer→character calibration (currently identity, fitting the face oval to the canvas); the override layer; anything on the Animator Pro side. The placeholder plate is a frozen face-oval polygon — it exists so the mouth has a face to read against, not to look good.