Audio-clocked playback for judging sync
Lip sync cannot be judged silently. extract.sh now pulls the audio track and writes manifest.json alongside the frames; the page reads the true extraction rate from it rather than assuming one, since a guessed fps desynchronises picture from sound - the one thing this view exists to show. Audio is the clock: frame = floor(currentTime * fps). A slow render loop drops frames instead of drifting, and half/quarter speed work via playbackRate with the picture following for free. Scrubbing, stepping and clicking a thumbnail all seek the audio too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
a082208ad5
commit
5ac919b9d6
7 changed files with 157 additions and 36 deletions
52
README.md
52
README.md
|
|
@ -24,7 +24,7 @@ python3 -m http.server 8777 # from this directory
|
|||
For real footage:
|
||||
|
||||
```sh
|
||||
./extract.sh /path/to/clip.mp4 24 # -> frames/0001.png …
|
||||
./extract.sh /path/to/clip.mov 12 # -> frames/*.png, audio.wav, manifest.json
|
||||
```
|
||||
|
||||
then **Load frames**. MediaPipe's wasm is fetched from jsdelivr on first use;
|
||||
|
|
@ -34,6 +34,14 @@ Frames are pre-extracted rather than decoded in the page because browser video
|
|||
seeking is approximate and `requestVideoFrameCallback` only delivers frames at
|
||||
playback speed — neither gives a deterministic per-frame pass.
|
||||
|
||||
`manifest.json` records the true extraction rate. The page reads it rather than
|
||||
assuming, because a guessed fps desynchronises audio from picture — and sync is
|
||||
the one thing this view exists to show.
|
||||
|
||||
**Audio is the playback clock**: `frame = floor(audio.currentTime * fps)`. A slow
|
||||
render loop therefore drops frames instead of drifting, and ½x / ¼x work by
|
||||
setting `playbackRate` with the picture following for free.
|
||||
|
||||
## Shooting for it
|
||||
|
||||
Near-frontal, good light, consistent scale, head reasonably still. Hold a neutral
|
||||
|
|
@ -49,18 +57,33 @@ is hand-drawn head plates, which this tool does not yet do.
|
|||
| Knob | What it does |
|
||||
| --- | --- |
|
||||
| vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. |
|
||||
| min hold | Minimum frames between keys. |
|
||||
| change gate | Mean vertex movement required before a new key is accepted. |
|
||||
| vel smoothing | Window on the velocity signal used to find extremes. |
|
||||
| anchor smoothing | Window on the four similarity parameters. Smooths the *transform*, never the contour. |
|
||||
| exposure | Grid that key frames snap onto. |
|
||||
| contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. |
|
||||
| anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform*. |
|
||||
| closed-mouth cut | Aperture below which the mouth interior is emitted as `hidden`. |
|
||||
| suggest tolerance | Max head movement before a new plate drawing is required. Affects **Suggest** only. |
|
||||
|
||||
Keys go on **velocity minima**, not distance thresholds: a threshold fires at the
|
||||
frame it was crossed — partway through a transition — so poses land mushy and
|
||||
late. The timeline shows the velocity curve, candidate minima, accepted keys
|
||||
(`f`) and their pre-snap extremes (`src`); a large `f`/`src` gap means min-hold
|
||||
and exposure are fighting.
|
||||
## Two kinds of sparseness
|
||||
|
||||
Sparseness has two unrelated causes, and conflating them was the original design
|
||||
error here. **Aesthetic** sparseness is set by the extraction rate — pick 12fps and
|
||||
you have already chosen your timing. **Labour** sparseness is a human drawing
|
||||
each one, and it binds only on the plate.
|
||||
|
||||
So the mouth keeps **every** frame: it is traced, and therefore free. In limited
|
||||
animation lip sync is routinely the densest element, on 1s, while heads hold on
|
||||
2s and 3s.
|
||||
|
||||
The frame strip is the editing surface for the other half: which frames need
|
||||
their own plate drawing. Everything starts kept; delete what you don't want.
|
||||
**Suggest** runs error-tolerance decimation over head pose as a starting point,
|
||||
then you hand-correct.
|
||||
|
||||
`contour avg` is a deliberate, bounded exception to "never smooth the contour" in
|
||||
`../docs/roto-puppet.md`. That rule held while keys were sparse, because sampling
|
||||
at velocity minima rejected detector noise for free. With a key on every frame it
|
||||
does not, so a radius shorter than the shortest articulation worth keeping is
|
||||
justified — at 12fps, articulation spans 3–6 frames and detector noise is
|
||||
per-frame, so ±1 separates them and ±3 starts eating speech.
|
||||
|
||||
## Tests
|
||||
|
||||
|
|
@ -79,8 +102,9 @@ eyeball.
|
|||
|
||||
## Not done yet
|
||||
|
||||
Eyes and irises; hand-drawn head plates and per-plate mouth slots; real
|
||||
Eyes and irises; hand-drawn head plates and per-plate mouth slots (the strip
|
||||
decides *which frames need one*, but you cannot yet supply the drawing); real
|
||||
performer→character calibration (currently identity, fitting the face oval to the
|
||||
canvas); the override layer; anything on the Animator Pro side. The placeholder
|
||||
plate is a frozen face-oval polygon — it exists so the mouth has a face to read
|
||||
canvas); the override layer; anything on the Animator Pro side. The plate is a
|
||||
face-oval polygon per kept frame — it exists so the mouth has a face to read
|
||||
against, not to look good.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue