Preserve source cadence and sample picture fps after analysis

This commit is contained in:
Olive Vaughn 2026-09-27 19:31:39 -04:00
parent 32683efccf
commit 06cf02db83
20 changed files with 303 additions and 87 deletions

View file

@ -42,7 +42,7 @@ wasm, which is fetched from a CDN on first use.
For real footage:
```sh
./extract.sh /path/to/clip.mov 12 # -> frames/*.png, audio.wav, manifest.json
./extract.sh /path/to/clip.mov # -> frames/*.png, audio.wav, manifest.json
```
then **Load frames**. MediaPipe's wasm is fetched from jsdelivr on first use;
@ -52,16 +52,16 @@ Frames are pre-extracted rather than decoded in the page because browser video
seeking is approximate and `requestVideoFrameCallback` only delivers frames at
playback speed — neither gives a deterministic per-frame pass.
`manifest.json` records the true extraction rate. The page reads it rather than
assuming, because a guessed fps desynchronises audio from picture — and sync is
the one thing this view exists to show.
`manifest.json` records the source rate. The extractor keeps every source frame;
the page reads that rate because a guessed fps desynchronises audio from picture.
Choosing a lower picture rate happens after analysis.
**Exposure** decides how often the picture gets a new drawing: rip at 24 and
render `on 2s` for 12, `on 3s` for 8. The dense track and the audio are
untouched, so it is a dropdown rather than a re-rip, and the export emits keys
only on the grid instead of the same pose twice. Everything rides the same grid
— mouth, eyes, teeth, plate — because a head cutting on the odd frames while the
mouth cuts on the even ones reads as two performances laid over each other.
**Picture fps** decides how often the finished roto gets a new pose. Analyze all
source frames, then sample those frozen poses at 12, 24 or the source rate while
keeping the original duration and audio. **Exposure** can hold a drawing across
more than one picture slot. The tracing editor chooses source frames for cel
references separately. Shared timing is the useful default for mouth, eyes,
teeth and plate so their changes read as one performance.
**Audio is the playback clock**: `frame = floor(audio.currentTime * fps)`. A slow
render loop therefore drops frames instead of drifting, and ½x / ¼x work by
@ -322,13 +322,13 @@ the tool a person made by hand; everything else regenerates. They are not in the
## Two kinds of sparseness
Sparseness has two unrelated causes, and conflating them was the original design
error here. **Aesthetic** sparseness is set by the extraction rate — pick 12fps and
you have already chosen your timing. **Labour** sparseness is a human drawing
each one, and it binds only on the plate.
error here. **Aesthetic** sparseness is chosen from the full analyzed source
track at rendering time. **Labour** sparseness is a human drawing each cel and
selecting which source frames to use as tracing references.
Aesthetic sparseness is the **exposure** control, not the extraction rate —
making it a render-time grid means auditioning 12 against 24 costs a dropdown
instead of a re-rip and a full re-detection.
Aesthetic sparseness is the **picture fps** control, with exposure available for
longer holds. Both happen after analysis, so auditioning 12 against 24 needs no
re-extraction or re-detection.
So the mouth keeps **every** frame: it is traced, and therefore free. In limited
animation lip sync is routinely the densest element, on 1s, while heads hold on