arthur/README.md
Your Name 5ac919b9d6 Audio-clocked playback for judging sync
Lip sync cannot be judged silently. extract.sh now pulls the audio track and
writes manifest.json alongside the frames; the page reads the true extraction
rate from it rather than assuming one, since a guessed fps desynchronises
picture from sound - the one thing this view exists to show.

Audio is the clock: frame = floor(currentTime * fps). A slow render loop drops
frames instead of drifting, and half/quarter speed work via playbackRate with
the picture following for free. Scrubbing, stepping and clicking a thumbnail
all seek the audio too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 14:53:09 -04:00

110 lines
4.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# roto
Video → **take file** builder for the Animator Pro rotoscope/puppet pipeline
described in `../docs/roto-puppet.md`.
This is the analysis and tuning half. It stabilises a face out of a clip, reduces
the lip contour to a handful of vertices, selects sparse keys on motion extremes,
and previews the result as flat indexed fills — so the *look and timing* can be
judged in seconds rather than through a minutes-long Animator Pro render. It
emits a `.take` file; nothing here touches Animator Pro.
All policy lives here. The take arrives at the renderer with its keys already
chosen.
## Run
```sh
python3 -m http.server 8777 # from this directory
# open http://127.0.0.1:8777
```
**Synthetic take** needs no video and exercises everything below detection.
For real footage:
```sh
./extract.sh /path/to/clip.mov 12 # -> frames/*.png, audio.wav, manifest.json
```
then **Load frames**. MediaPipe's wasm is fetched from jsdelivr on first use;
`face_landmarker.task` is local.
Frames are pre-extracted rather than decoded in the page because browser video
seeking is approximate and `requestVideoFrameCallback` only delivers frames at
playback speed — neither gives a deterministic per-frame pass.
`manifest.json` records the true extraction rate. The page reads it rather than
assuming, because a guessed fps desynchronises audio from picture — and sync is
the one thing this view exists to show.
**Audio is the playback clock**: `frame = floor(audio.currentTime * fps)`. A slow
render loop therefore drops frames instead of drifting, and ½x / ¼x work by
setting `playbackRate` with the picture following for free.
## Shooting for it
Near-frontal, good light, consistent scale, head reasonably still. Hold a neutral
closed mouth for a second at the top of the take: that frame is picked
automatically as the neutral and drives calibration and the placeholder plate.
Out-of-plane head rotation cannot be stabilised away by a 2D similarity
transform — the `residual` readout rises when it happens. The pipeline's answer
is hand-drawn head plates, which this tool does not yet do.
## Knobs
| Knob | What it does |
| --- | --- |
| vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. |
| contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. |
| anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform*. |
| closed-mouth cut | Aperture below which the mouth interior is emitted as `hidden`. |
| suggest tolerance | Max head movement before a new plate drawing is required. Affects **Suggest** only. |
## Two kinds of sparseness
Sparseness has two unrelated causes, and conflating them was the original design
error here. **Aesthetic** sparseness is set by the extraction rate — pick 12fps and
you have already chosen your timing. **Labour** sparseness is a human drawing
each one, and it binds only on the plate.
So the mouth keeps **every** frame: it is traced, and therefore free. In limited
animation lip sync is routinely the densest element, on 1s, while heads hold on
2s and 3s.
The frame strip is the editing surface for the other half: which frames need
their own plate drawing. Everything starts kept; delete what you don't want.
**Suggest** runs error-tolerance decimation over head pose as a starting point,
then you hand-correct.
`contour avg` is a deliberate, bounded exception to "never smooth the contour" in
`../docs/roto-puppet.md`. That rule held while keys were sparse, because sampling
at velocity minima rejected detector noise for free. With a key on every frame it
does not, so a radius shorter than the shortest articulation worth keeping is
justified — at 12fps, articulation spans 3–6 frames and detector noise is
per-frame, so ±1 separates them and ±3 starts eating speech.
## Tests
```sh
chromium --headless --virtual-time-budget=8000 --dump-dom \
http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+'
```
Or open `selftest.html`. 29 assertions over the stages below detection.
The ring-simplicity check is the load-bearing one. Because `hold` parts *cut*
between poses instead of interpolating, a ring whose vertex order is wrong
self-intersects and renders as blocks meeting at corners — and it is invisible at
odd `verts/2` and obvious at even, so it needs an assertion rather than an
eyeball.
## Not done yet
Eyes and irises; hand-drawn head plates and per-plate mouth slots (the strip
decides *which frames need one*, but you cannot yet supply the drawing); real
performer→character calibration (currently identity, fitting the face oval to the
canvas); the override layer; anything on the Animator Pro side. The plate is a
face-oval polygon per kept frame — it exists so the mouth has a face to read
against, not to look good.