Frame removal for plates; mouth keeps every frame
Sparseness was being applied for two different reasons at once. Aesthetic sparseness is set by the extraction rate; labour sparseness only binds on the plate, because a human draws each one. The mouth is traced and therefore free, and in limited animation lip sync is routinely the densest element - on 1s while heads hold on 2s and 3s. So: the mouth gets a key on every frame, and the frame strip is now the editing surface for deciding which frames need their own plate drawing. All frames start kept; delete the ones you don't want. - strip of face-cropped thumbnails, keep/drop per frame, keyboard driven - worksheet panel lists the drawings needed and the range each one holds - Suggest runs error-tolerance decimation on head pose as a starting point - export writes sparse plate keys + dense mouth keys, with a hold manifest - smoothContours: bounded exception to "never smooth the contour", which held only while keys were sparse enough to reject detector noise by sampling - averages are now a RADIUS in frames: 0 is off, 1 is +-1 - GPU delegate falls back to CPU instead of failing - #synth / #frames autorun for headless smoke tests Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
7a6bdea157
commit
a082208ad5
7 changed files with 489 additions and 265 deletions
|
|
@ -10,15 +10,19 @@ const pick = (lm, idx) => idx.map((i) => ({ x: lm[i].x, y: lm[i].y }));
|
|||
// Stage 1-3: fit the rigid transform per frame, smooth its parameters, then map
|
||||
// every contour through it into the reference frame. The result is head-local:
|
||||
// translation, roll and depth-scale of the head are gone.
|
||||
export function stabilize(dense, smoothWin) {
|
||||
export function stabilize(dense, smoothRadius) {
|
||||
const rigid = dense.map((f) => pick(f, RIGID));
|
||||
const ref = procrustesMean(rigid);
|
||||
const raw = rigid.map((r) => fitSimilarity(r, ref));
|
||||
const tfs = smoothTransforms(raw, smoothWin);
|
||||
const tfs = smoothTransforms(raw, smoothRadius);
|
||||
|
||||
return {
|
||||
ref,
|
||||
transforms: tfs,
|
||||
// Rigid landmarks in IMAGE space: the head-pose signal. Frame removal is
|
||||
// decided from head motion, not from the mouth, so this has to survive the
|
||||
// fit rather than being consumed by it.
|
||||
rigid,
|
||||
// Residual rises with out-of-plane rotation, which no 2D similarity can
|
||||
// remove. High values mean this section wants a different head plate.
|
||||
residual: tfs.map((tf, i) => fitResidual(tf, rigid[i], ref)),
|
||||
|
|
@ -106,3 +110,61 @@ export function activeKey(keys, f) {
|
|||
for (const k of keys) { if (k.f <= f) hit = k; else break; }
|
||||
return hit;
|
||||
}
|
||||
|
||||
// Temporal smoothing of a contour, per vertex, across time.
|
||||
//
|
||||
// docs/roto-puppet.md says to smooth the transform and never the contour. That
|
||||
// was correct while keys were sparse: sampling at velocity minima rejected
|
||||
// per-frame detector noise for free. With a key on every frame the noise is
|
||||
// visible as a shimmer along the lip edge, so a bounded exception applies -
|
||||
// the window must stay SHORTER than the shortest articulation worth keeping.
|
||||
// At 12fps, mouth movement spans 3-6 frames and detector noise is per-frame, so
|
||||
// a radius of 1 separates them and a radius of 3 would start eating speech.
|
||||
//
|
||||
// `radius` in frames either side: 0 off, 1 = 3-frame average, 2 = 5-frame.
|
||||
export function smoothContours(rings, radius) {
|
||||
if (radius <= 0) return rings;
|
||||
const half = Math.floor(radius), N = rings.length, V = rings[0].length;
|
||||
const out = [];
|
||||
for (let t = 0; t < N; t++) {
|
||||
const frame = [];
|
||||
for (let v = 0; v < V; v++) {
|
||||
let sx = 0, sy = 0, c = 0;
|
||||
for (let j = t - half; j <= t + half; j++) {
|
||||
const k = Math.min(N - 1, Math.max(0, j));
|
||||
sx += rings[k][v].x; sy += rings[k][v].y; c++;
|
||||
}
|
||||
frame.push({ x: sx / c, y: sy / c });
|
||||
}
|
||||
out.push(frame);
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
// Which frames need their own PLATE drawing.
|
||||
//
|
||||
// This is frame removal, not keyframe extraction: every frame is a candidate and
|
||||
// the question is which can be dropped. Walk forward holding the current drawing
|
||||
// until the head has moved further than `tol` from it, then a new drawing is
|
||||
// required. The cost being managed is an artist drawing a head, which is why the
|
||||
// signal is head pose and not the mouth - the mouth is traced and free.
|
||||
export function suggestPlateFrames(rigid, tol) {
|
||||
const dist = (a, b) => {
|
||||
let m = 0;
|
||||
for (let i = 0; i < a.length; i++) m = Math.max(m, Math.hypot(a[i].x - b[i].x, a[i].y - b[i].y));
|
||||
return m;
|
||||
};
|
||||
const keep = [0];
|
||||
let anchor = 0;
|
||||
for (let f = 1; f < rigid.length; f++) {
|
||||
if (dist(rigid[f], rigid[anchor]) > tol) { keep.push(f); anchor = f; }
|
||||
}
|
||||
return keep;
|
||||
}
|
||||
|
||||
// Nearest kept frame at or before f - the plate that is on screen.
|
||||
export function heldFrame(kept, f) {
|
||||
let hit = kept[0];
|
||||
for (const k of kept) { if (k <= f) hit = k; else break; }
|
||||
return hit;
|
||||
}
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue