Preserve source cadence and sample picture fps after analysis
This commit is contained in:
parent
32683efccf
commit
06cf02db83
20 changed files with 303 additions and 87 deletions
32
README.md
32
README.md
|
|
@ -42,7 +42,7 @@ wasm, which is fetched from a CDN on first use.
|
|||
For real footage:
|
||||
|
||||
```sh
|
||||
./extract.sh /path/to/clip.mov 12 # -> frames/*.png, audio.wav, manifest.json
|
||||
./extract.sh /path/to/clip.mov # -> frames/*.png, audio.wav, manifest.json
|
||||
```
|
||||
|
||||
then **Load frames**. MediaPipe's wasm is fetched from jsdelivr on first use;
|
||||
|
|
@ -52,16 +52,16 @@ Frames are pre-extracted rather than decoded in the page because browser video
|
|||
seeking is approximate and `requestVideoFrameCallback` only delivers frames at
|
||||
playback speed — neither gives a deterministic per-frame pass.
|
||||
|
||||
`manifest.json` records the true extraction rate. The page reads it rather than
|
||||
assuming, because a guessed fps desynchronises audio from picture — and sync is
|
||||
the one thing this view exists to show.
|
||||
`manifest.json` records the source rate. The extractor keeps every source frame;
|
||||
the page reads that rate because a guessed fps desynchronises audio from picture.
|
||||
Choosing a lower picture rate happens after analysis.
|
||||
|
||||
**Exposure** decides how often the picture gets a new drawing: rip at 24 and
|
||||
render `on 2s` for 12, `on 3s` for 8. The dense track and the audio are
|
||||
untouched, so it is a dropdown rather than a re-rip, and the export emits keys
|
||||
only on the grid instead of the same pose twice. Everything rides the same grid
|
||||
— mouth, eyes, teeth, plate — because a head cutting on the odd frames while the
|
||||
mouth cuts on the even ones reads as two performances laid over each other.
|
||||
**Picture fps** decides how often the finished roto gets a new pose. Analyze all
|
||||
source frames, then sample those frozen poses at 12, 24 or the source rate while
|
||||
keeping the original duration and audio. **Exposure** can hold a drawing across
|
||||
more than one picture slot. The tracing editor chooses source frames for cel
|
||||
references separately. Shared timing is the useful default for mouth, eyes,
|
||||
teeth and plate so their changes read as one performance.
|
||||
|
||||
**Audio is the playback clock**: `frame = floor(audio.currentTime * fps)`. A slow
|
||||
render loop therefore drops frames instead of drifting, and ½x / ¼x work by
|
||||
|
|
@ -322,13 +322,13 @@ the tool a person made by hand; everything else regenerates. They are not in the
|
|||
## Two kinds of sparseness
|
||||
|
||||
Sparseness has two unrelated causes, and conflating them was the original design
|
||||
error here. **Aesthetic** sparseness is set by the extraction rate — pick 12fps and
|
||||
you have already chosen your timing. **Labour** sparseness is a human drawing
|
||||
each one, and it binds only on the plate.
|
||||
error here. **Aesthetic** sparseness is chosen from the full analyzed source
|
||||
track at rendering time. **Labour** sparseness is a human drawing each cel and
|
||||
selecting which source frames to use as tracing references.
|
||||
|
||||
Aesthetic sparseness is the **exposure** control, not the extraction rate —
|
||||
making it a render-time grid means auditioning 12 against 24 costs a dropdown
|
||||
instead of a re-rip and a full re-detection.
|
||||
Aesthetic sparseness is the **picture fps** control, with exposure available for
|
||||
longer holds. Both happen after analysis, so auditioning 12 against 24 needs no
|
||||
re-extraction or re-detection.
|
||||
|
||||
So the mouth keeps **every** frame: it is traced, and therefore free. In limited
|
||||
animation lip sync is routinely the densest element, on 1s, while heads hold on
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue