roto: video -> take file builder with interactive tuning
Analysis half of the pipeline in docs/roto-puppet.md. Stabilises a face out of a clip via a similarity fit on rigid landmarks, reduces the lip contour to a fixed vertex budget, selects sparse keys on velocity minima, and previews the result as flat indexed fills so timing can be judged without an Animator Pro render. - landmarks.js ordered lip/oval rings; slot position is vertex identity - mathutil.js closed-form 2D similarity, Procrustes mean, transform smoothing - pipeline.js stabilise -> subsample -> key-select - raster.js indexed scanline fill, no antialiasing - take.js take-file writer - selftest.js 29 assertions, incl. ring simplicity at every vertex budget Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
commit
7a6bdea157
13 changed files with 1277 additions and 0 deletions
3
.gitignore
vendored
Normal file
3
.gitignore
vendored
Normal file
|
|
@ -0,0 +1,3 @@
|
|||
frames/
|
||||
*.task
|
||||
*.take
|
||||
86
README.md
Normal file
86
README.md
Normal file
|
|
@ -0,0 +1,86 @@
|
|||
# roto
|
||||
|
||||
Video → **take file** builder for the Animator Pro rotoscope/puppet pipeline
|
||||
described in `../docs/roto-puppet.md`.
|
||||
|
||||
This is the analysis and tuning half. It stabilises a face out of a clip, reduces
|
||||
the lip contour to a handful of vertices, selects sparse keys on motion extremes,
|
||||
and previews the result as flat indexed fills — so the *look and timing* can be
|
||||
judged in seconds rather than through a minutes-long Animator Pro render. It
|
||||
emits a `.take` file; nothing here touches Animator Pro.
|
||||
|
||||
All policy lives here. The take arrives at the renderer with its keys already
|
||||
chosen.
|
||||
|
||||
## Run
|
||||
|
||||
```sh
|
||||
python3 -m http.server 8777 # from this directory
|
||||
# open http://127.0.0.1:8777
|
||||
```
|
||||
|
||||
**Synthetic take** needs no video and exercises everything below detection.
|
||||
|
||||
For real footage:
|
||||
|
||||
```sh
|
||||
./extract.sh /path/to/clip.mp4 24 # -> frames/0001.png …
|
||||
```
|
||||
|
||||
then **Load frames**. MediaPipe's wasm is fetched from jsdelivr on first use;
|
||||
`face_landmarker.task` is local.
|
||||
|
||||
Frames are pre-extracted rather than decoded in the page because browser video
|
||||
seeking is approximate and `requestVideoFrameCallback` only delivers frames at
|
||||
playback speed — neither gives a deterministic per-frame pass.
|
||||
|
||||
## Shooting for it
|
||||
|
||||
Near-frontal, good light, consistent scale, head reasonably still. Hold a neutral
|
||||
closed mouth for a second at the top of the take: that frame is picked
|
||||
automatically as the neutral and drives calibration and the placeholder plate.
|
||||
|
||||
Out-of-plane head rotation cannot be stabilised away by a 2D similarity
|
||||
transform — the `residual` readout rises when it happens. The pipeline's answer
|
||||
is hand-drawn head plates, which this tool does not yet do.
|
||||
|
||||
## Knobs
|
||||
|
||||
| Knob | What it does |
|
||||
| --- | --- |
|
||||
| vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. |
|
||||
| min hold | Minimum frames between keys. |
|
||||
| change gate | Mean vertex movement required before a new key is accepted. |
|
||||
| vel smoothing | Window on the velocity signal used to find extremes. |
|
||||
| anchor smoothing | Window on the four similarity parameters. Smooths the *transform*, never the contour. |
|
||||
| exposure | Grid that key frames snap onto. |
|
||||
| closed-mouth cut | Aperture below which the mouth interior is emitted as `hidden`. |
|
||||
|
||||
Keys go on **velocity minima**, not distance thresholds: a threshold fires at the
|
||||
frame it was crossed — partway through a transition — so poses land mushy and
|
||||
late. The timeline shows the velocity curve, candidate minima, accepted keys
|
||||
(`f`) and their pre-snap extremes (`src`); a large `f`/`src` gap means min-hold
|
||||
and exposure are fighting.
|
||||
|
||||
## Tests
|
||||
|
||||
```sh
|
||||
chromium --headless --virtual-time-budget=8000 --dump-dom \
|
||||
http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+'
|
||||
```
|
||||
|
||||
Or open `selftest.html`. 29 assertions over the stages below detection.
|
||||
|
||||
The ring-simplicity check is the load-bearing one. Because `hold` parts *cut*
|
||||
between poses instead of interpolating, a ring whose vertex order is wrong
|
||||
self-intersects and renders as blocks meeting at corners — and it is invisible at
|
||||
odd `verts/2` and obvious at even, so it needs an assertion rather than an
|
||||
eyeball.
|
||||
|
||||
## Not done yet
|
||||
|
||||
Eyes and irises; hand-drawn head plates and per-plate mouth slots; real
|
||||
performer→character calibration (currently identity, fitting the face oval to the
|
||||
canvas); the override layer; anything on the Animator Pro side. The placeholder
|
||||
plate is a frozen face-oval polygon — it exists so the mouth has a face to read
|
||||
against, not to look good.
|
||||
17
extract.sh
Executable file
17
extract.sh
Executable file
|
|
@ -0,0 +1,17 @@
|
|||
#!/usr/bin/env bash
|
||||
# Extract a clip to a PNG sequence for the take builder.
|
||||
#
|
||||
# Frames are pre-extracted rather than decoded in the page on purpose: browser
|
||||
# video seeking by currentTime is approximate and requestVideoFrameCallback only
|
||||
# delivers frames at playback speed, so neither gives a deterministic per-frame
|
||||
# pass. A PNG sequence is exact, instantly seekable, and reproducible.
|
||||
set -euo pipefail
|
||||
|
||||
src="${1:?usage: ./extract.sh CLIP [FPS] [OUTDIR]}"
|
||||
fps="${2:-24}"
|
||||
out="${3:-frames}"
|
||||
|
||||
rm -rf "$out"
|
||||
mkdir -p "$out"
|
||||
ffmpeg -hide_banner -loglevel warning -i "$src" -vf "fps=$fps" "$out/%04d.png"
|
||||
echo "$(ls -1 "$out" | wc -l) frames at ${fps}fps -> $out/"
|
||||
123
index.html
Normal file
123
index.html
Normal file
|
|
@ -0,0 +1,123 @@
|
|||
<!doctype html>
|
||||
<html lang="en">
|
||||
<head>
|
||||
<meta charset="utf-8">
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||||
<title>roto — take builder</title>
|
||||
<style>
|
||||
:root {
|
||||
--bg: #0b0d13; --panel: #12151f; --line: #232836; --fg: #e8eaf0;
|
||||
--dim: #8891a5; --accent: #fbbf24; --ok: #4ade80; --err: #f87171;
|
||||
}
|
||||
* { box-sizing: border-box; }
|
||||
body { margin: 0; background: var(--bg); color: var(--fg);
|
||||
font: 13px/1.5 ui-monospace, SFMono-Regular, Menlo, monospace; }
|
||||
header { padding: 14px 18px; border-bottom: 1px solid var(--line);
|
||||
display: flex; gap: 14px; align-items: baseline; flex-wrap: wrap; }
|
||||
h1 { font-size: 14px; margin: 0; letter-spacing: .04em; }
|
||||
h1 span { color: var(--dim); font-weight: 400; }
|
||||
main { padding: 18px; display: flex; flex-direction: column; gap: 18px; }
|
||||
.row { display: flex; gap: 18px; flex-wrap: wrap; }
|
||||
.panel { background: var(--panel); border: 1px solid var(--line); border-radius: 4px; padding: 12px; }
|
||||
.panel h2 { font-size: 11px; text-transform: uppercase; letter-spacing: .08em;
|
||||
color: var(--dim); margin: 0 0 8px; font-weight: 500; }
|
||||
canvas { display: block; image-rendering: pixelated; max-width: 100%; border-radius: 2px; }
|
||||
#cv-timeline { width: 100%; height: 90px; }
|
||||
button, input[type=text] {
|
||||
font: inherit; background: #1c2130; color: var(--fg);
|
||||
border: 1px solid var(--line); border-radius: 3px; padding: 5px 11px; cursor: pointer; }
|
||||
button:hover { border-color: var(--accent); }
|
||||
input[type=text] { cursor: text; }
|
||||
label.ctl { display: grid; grid-template-columns: 116px 1fr 44px;
|
||||
gap: 10px; align-items: center; margin-bottom: 6px; }
|
||||
label.ctl span:first-child { color: var(--dim); }
|
||||
label.ctl output { text-align: right; color: var(--accent); }
|
||||
input[type=range] { width: 100%; accent-color: var(--accent); }
|
||||
#status { color: var(--dim); }
|
||||
#status.ok { color: var(--ok); } #status.err { color: var(--err); } #status.warn { color: var(--accent); }
|
||||
#readout { color: var(--dim); font-size: 12px; }
|
||||
#sheet { display: flex; flex-wrap: wrap; gap: 8px; }
|
||||
#sheet .cell { display: flex; flex-direction: column; gap: 3px; cursor: pointer; }
|
||||
#sheet .cell span { color: var(--dim); font-size: 11px; }
|
||||
#sheet .cell:hover span { color: var(--accent); }
|
||||
#palette { display: flex; gap: 12px; flex-wrap: wrap; }
|
||||
.sw { display: flex; align-items: center; gap: 5px; color: var(--dim); font-size: 11px; }
|
||||
.sw input { width: 26px; height: 20px; padding: 0; border: 1px solid var(--line); background: none; }
|
||||
.legend { color: var(--dim); font-size: 11px; margin-top: 6px; }
|
||||
.legend b { font-weight: 400; }
|
||||
.k { color: var(--accent); } .s { color: var(--err); } .c { color: #5b6478; }
|
||||
</style>
|
||||
</head>
|
||||
<body>
|
||||
<header>
|
||||
<h1>roto <span>— video → take builder</span></h1>
|
||||
<button id="btn-synth">Synthetic take</button>
|
||||
<input type="text" id="framedir" value="frames" size="8" title="frame directory">
|
||||
<button id="btn-frames">Load frames</button>
|
||||
<button id="btn-play">Play</button>
|
||||
<input type="text" id="takename" value="line_01" size="10" title="take name">
|
||||
<button id="btn-export">Export .take</button>
|
||||
<span id="status"></span>
|
||||
</header>
|
||||
|
||||
<main>
|
||||
<div class="row">
|
||||
<div class="panel">
|
||||
<h2>source + landmarks</h2>
|
||||
<canvas id="cv-source"></canvas>
|
||||
<div class="legend">outer lip <b style="color:#4ade80">—</b> · inner lip <b style="color:#f87171">—</b></div>
|
||||
</div>
|
||||
<div class="panel">
|
||||
<h2>stabilised (head-local)</h2>
|
||||
<canvas id="cv-stab"></canvas>
|
||||
<div class="legend">live contour · <b class="k">—</b> active key pose · should sit still except the mouth</div>
|
||||
</div>
|
||||
<div class="panel">
|
||||
<h2>flat render — 320×200 indexed</h2>
|
||||
<canvas id="cv-render"></canvas>
|
||||
<div class="legend" id="framelabel"></div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div class="row">
|
||||
<div class="panel" style="flex:1 1 380px">
|
||||
<h2>knobs</h2>
|
||||
<label class="ctl"><span>vertices</span><input type="range" id="verts" min="4" max="16" step="2" value="8"><output id="vertsv"></output></label>
|
||||
<label class="ctl"><span>min hold</span><input type="range" id="minHold" min="1" max="8" value="2"><output id="minHoldv"></output></label>
|
||||
<label class="ctl"><span>change gate</span><input type="range" id="distThresh" min="0" max="40" value="6"><output id="distThreshv"></output></label>
|
||||
<label class="ctl"><span>vel smoothing</span><input type="range" id="velSmooth" min="1" max="11" step="2" value="3"><output id="velSmoothv"></output></label>
|
||||
<label class="ctl"><span>anchor smoothing</span><input type="range" id="smoothWin" min="1" max="21" step="2" value="5"><output id="smoothWinv"></output></label>
|
||||
<label class="ctl"><span>exposure</span><input type="range" id="exposure" min="1" max="4" value="2"><output id="exposurev"></output></label>
|
||||
<label class="ctl"><span>closed-mouth cut</span><input type="range" id="apertureThresh" min="0" max="400" value="120"><output id="apertureThreshv"></output></label>
|
||||
<div class="legend" id="readout"></div>
|
||||
</div>
|
||||
<div class="panel" style="flex:1 1 300px">
|
||||
<h2>palette</h2>
|
||||
<div id="palette"></div>
|
||||
<div class="legend" style="margin-top:12px">
|
||||
Flat indexed fills, no antialiasing — the rasteriser writes palette
|
||||
indices, the way <code>csd_render_poly</code> does.
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div class="panel">
|
||||
<h2>timeline</h2>
|
||||
<canvas id="cv-timeline"></canvas>
|
||||
<input type="range" id="scrub" min="0" max="0" value="0" style="width:100%;margin-top:8px">
|
||||
<div class="legend">
|
||||
velocity curve · <b class="c">|</b> candidate minima ·
|
||||
<b class="k">|</b> key as rendered (<code>f</code>) ·
|
||||
<b class="s">|</b> true extreme (<code>src</code>) — a large gap means min-hold and exposure are fighting
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<div class="panel">
|
||||
<h2>selected poses — the vocabulary the take actually contains</h2>
|
||||
<div id="sheet"></div>
|
||||
</div>
|
||||
</main>
|
||||
|
||||
<script type="module" src="./js/app.js"></script>
|
||||
</body>
|
||||
</html>
|
||||
403
js/app.js
Normal file
403
js/app.js
Normal file
|
|
@ -0,0 +1,403 @@
|
|||
import { FaceLandmarker, FilesetResolver } from 'https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@1.0.1/vision_bundle.mjs';
|
||||
import { LIPS_OUTER, LIPS_INNER, FACE_OVAL } from './landmarks.js';
|
||||
import { stabilize, toRasterRing, selectKeys, activeKey } from './pipeline.js';
|
||||
import { IndexedRaster } from './raster.js';
|
||||
import { writeTake } from './take.js';
|
||||
import { synthDense } from './synth.js';
|
||||
|
||||
const RW = 320, RH = 200, ZOOM = 2;
|
||||
|
||||
const PALETTE = [
|
||||
{ name: 'bg', hex: '#12141c' },
|
||||
{ name: 'skin_base', hex: '#b07a5a' },
|
||||
{ name: 'skin_dark', hex: '#7a4f3a' },
|
||||
{ name: 'mouth_dark', hex: '#24161a' },
|
||||
{ name: 'skin_lite', hex: '#d9a884' },
|
||||
];
|
||||
const IDX = { bg: 0, base: 1, dark: 2, mouth: 3, lite: 4 };
|
||||
|
||||
const state = {
|
||||
dense: null, images: [], stab: null, xform: null,
|
||||
plate: null, neutral: 0, keysOuter: null, keysInner: null,
|
||||
frame: 0, playing: false, source: 'none',
|
||||
};
|
||||
|
||||
const el = (id) => document.getElementById(id);
|
||||
const opts = () => ({
|
||||
verts: +el('verts').value,
|
||||
minHold: +el('minHold').value,
|
||||
distThresh: +el('distThresh').value / 10,
|
||||
velSmooth: +el('velSmooth').value,
|
||||
smoothWin: +el('smoothWin').value,
|
||||
exposure: +el('exposure').value,
|
||||
apertureThresh: +el('apertureThresh').value / 1000,
|
||||
});
|
||||
|
||||
function status(msg, kind = '') {
|
||||
const s = el('status');
|
||||
s.textContent = msg;
|
||||
s.className = kind;
|
||||
}
|
||||
|
||||
/* ---------- frame loading ---------- */
|
||||
|
||||
function loadImage(src) {
|
||||
return new Promise((res) => {
|
||||
const im = new Image();
|
||||
im.onload = () => res(im);
|
||||
im.onerror = () => res(null);
|
||||
im.src = src;
|
||||
});
|
||||
}
|
||||
|
||||
async function loadFrameSequence() {
|
||||
const dir = el('framedir').value.replace(/\/$/, '');
|
||||
const imgs = [];
|
||||
for (let i = 1; i <= 900; i++) {
|
||||
const name = `${dir}/${String(i).padStart(4, '0')}.png`;
|
||||
const im = await loadImage(name);
|
||||
if (!im) break;
|
||||
imgs.push(im);
|
||||
if (i % 10 === 0) status(`loading frames… ${i}`);
|
||||
}
|
||||
return imgs;
|
||||
}
|
||||
|
||||
/* ---------- detection ---------- */
|
||||
|
||||
let landmarker = null;
|
||||
|
||||
async function initLandmarker() {
|
||||
if (landmarker) return landmarker;
|
||||
status('loading MediaPipe wasm…');
|
||||
const fileset = await FilesetResolver.forVisionTasks(
|
||||
'https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@1.0.1/wasm');
|
||||
landmarker = await FaceLandmarker.createFromOptions(fileset, {
|
||||
baseOptions: { modelAssetPath: './face_landmarker.task', delegate: 'GPU' },
|
||||
runningMode: 'IMAGE',
|
||||
numFaces: 1,
|
||||
});
|
||||
return landmarker;
|
||||
}
|
||||
|
||||
async function detectAll(images) {
|
||||
const lm = await initLandmarker();
|
||||
const cv = document.createElement('canvas');
|
||||
const dense = [];
|
||||
const missing = [];
|
||||
for (let i = 0; i < images.length; i++) {
|
||||
const im = images[i];
|
||||
cv.width = im.naturalWidth; cv.height = im.naturalHeight;
|
||||
cv.getContext('2d').drawImage(im, 0, 0);
|
||||
const out = lm.detect(cv);
|
||||
if (out.faceLandmarks && out.faceLandmarks.length) {
|
||||
dense.push(out.faceLandmarks[0]);
|
||||
} else {
|
||||
// Hold the previous frame rather than dropping it, so frame indices stay
|
||||
// aligned with the source sequence. A gap is reported, not hidden.
|
||||
missing.push(i);
|
||||
dense.push(dense.length ? dense[dense.length - 1] : null);
|
||||
}
|
||||
if (i % 5 === 0) status(`detecting… ${i + 1}/${images.length}`);
|
||||
}
|
||||
if (dense[0] === null) throw new Error('no face found in the first frame');
|
||||
return { dense, missing };
|
||||
}
|
||||
|
||||
/* ---------- calibration ---------- */
|
||||
|
||||
// v1 calibration: fit the reference face oval's bounding box to a target box on
|
||||
// the character raster. Identity retargeting - it makes any clip frame sensibly,
|
||||
// but a real project replaces this with a measured performer->character map.
|
||||
function makeXform(stab) {
|
||||
const oval = stab.oval[state.neutral];
|
||||
let x0 = Infinity, y0 = Infinity, x1 = -Infinity, y1 = -Infinity;
|
||||
for (const p of oval) {
|
||||
x0 = Math.min(x0, p.x); y0 = Math.min(y0, p.y);
|
||||
x1 = Math.max(x1, p.x); y1 = Math.max(y1, p.y);
|
||||
}
|
||||
const targetH = RH * 0.80;
|
||||
const s = targetH / (y1 - y0);
|
||||
const cx = (x0 + x1) / 2, cy = (y0 + y1) / 2;
|
||||
return (p) => ({ x: (p.x - cx) * s + RW / 2, y: (p.y - cy) * s + RH / 2 });
|
||||
}
|
||||
|
||||
/* ---------- build ---------- */
|
||||
|
||||
function rebuild() {
|
||||
if (!state.dense) return;
|
||||
const o = opts();
|
||||
|
||||
state.stab = stabilize(state.dense, o.smoothWin);
|
||||
|
||||
// Neutral frame = most closed mouth in the first quarter of the shot, which is
|
||||
// where the performer is asked to hold a neutral closed mouth.
|
||||
const ap = state.stab.aperture;
|
||||
const head = Math.max(1, Math.floor(ap.length / 4));
|
||||
let best = 0;
|
||||
for (let i = 0; i < head; i++) if (ap[i] < ap[best]) best = i;
|
||||
state.neutral = best;
|
||||
|
||||
state.xform = makeXform(state.stab);
|
||||
state.plate = state.stab.oval[state.neutral].map(state.xform);
|
||||
|
||||
const outer = state.stab.outer.map((r) => toRasterRing(r, LIPS_OUTER, o.verts, state.xform));
|
||||
const inner = state.stab.inner.map((r) => toRasterRing(r, LIPS_INNER, o.verts, state.xform));
|
||||
|
||||
state.shapesOuter = outer;
|
||||
state.shapesInner = inner;
|
||||
state.keysOuter = selectKeys(outer, o);
|
||||
// The interior is a child of the lip silhouette: it keys on exactly its
|
||||
// parent's frames, never independently, or it swims.
|
||||
state.keysInner = {
|
||||
...state.keysOuter,
|
||||
keys: state.keysOuter.keys.map((k) => ({ ...k })),
|
||||
};
|
||||
|
||||
const apMax = Math.max(...ap);
|
||||
state.hidden = state.keysOuter.keys.map((k) => ap[k.src] / apMax < o.apertureThresh);
|
||||
|
||||
el('readout').textContent =
|
||||
`${state.dense.length} frames · ${state.keysOuter.candidates.length} candidates → ` +
|
||||
`${state.keysOuter.keys.length} keys · ${o.verts} verts · neutral f${state.neutral} · ` +
|
||||
`residual ${(state.stab.residual.reduce((a, b) => a + b, 0) / state.dense.length).toFixed(4)}`;
|
||||
|
||||
drawTimeline();
|
||||
drawContactSheet();
|
||||
draw();
|
||||
}
|
||||
|
||||
/* ---------- drawing ---------- */
|
||||
|
||||
function renderFrame(f) {
|
||||
const r = new IndexedRaster(RW, RH);
|
||||
r.clear(IDX.bg);
|
||||
if (state.plate) r.fillPoly(state.plate, IDX.base);
|
||||
|
||||
const ko = activeKey(state.keysOuter.keys, f);
|
||||
const ki = activeKey(state.keysInner.keys, f);
|
||||
if (ko) r.fillPoly(state.shapesOuter[ko.src], IDX.dark);
|
||||
if (ki) {
|
||||
const slot = state.keysOuter.keys.indexOf(ki);
|
||||
if (!state.hidden[slot]) r.fillPoly(state.shapesInner[ki.src], IDX.mouth);
|
||||
}
|
||||
return r;
|
||||
}
|
||||
|
||||
function blit(canvas, raster, zoom) {
|
||||
canvas.width = RW * zoom; canvas.height = RH * zoom;
|
||||
canvas.getContext('2d').putImageData(raster.toImageData(PALETTE.map((p) => p.hex), zoom), 0, 0);
|
||||
}
|
||||
|
||||
function draw() {
|
||||
const f = state.frame;
|
||||
el('framelabel').textContent =
|
||||
`f ${String(f).padStart(3)} / ${state.dense.length - 1} ` +
|
||||
`key ${activeKey(state.keysOuter.keys, f).f}`;
|
||||
|
||||
// pane 1: source with the raw contour overlaid
|
||||
const c1 = el('cv-source'), g1 = c1.getContext('2d');
|
||||
c1.width = RW * ZOOM; c1.height = RH * ZOOM;
|
||||
g1.fillStyle = '#000'; g1.fillRect(0, 0, c1.width, c1.height);
|
||||
const im = state.images[f];
|
||||
if (im) {
|
||||
const s = Math.min(c1.width / im.naturalWidth, c1.height / im.naturalHeight);
|
||||
const w = im.naturalWidth * s, h = im.naturalHeight * s;
|
||||
g1.drawImage(im, (c1.width - w) / 2, (c1.height - h) / 2, w, h);
|
||||
g1.save();
|
||||
g1.translate((c1.width - w) / 2, (c1.height - h) / 2);
|
||||
strokeRing(g1, LIPS_OUTER.map((i) => state.dense[f][i]), w, h, '#4ade80');
|
||||
strokeRing(g1, LIPS_INNER.map((i) => state.dense[f][i]), w, h, '#f87171');
|
||||
g1.restore();
|
||||
} else {
|
||||
g1.fillStyle = '#555'; g1.font = '13px system-ui';
|
||||
g1.fillText('synthetic — no source frames', 14, 24);
|
||||
strokeRing(g1, LIPS_OUTER.map((i) => state.dense[f][i]), c1.width, c1.height, '#4ade80');
|
||||
strokeRing(g1, LIPS_INNER.map((i) => state.dense[f][i]), c1.width, c1.height, '#f87171');
|
||||
}
|
||||
|
||||
// pane 2: stabilised contour against a fixed reference cross.
|
||||
// If stabilisation works, this contour stays put except for mouth motion.
|
||||
const c2 = el('cv-stab'), g2 = c2.getContext('2d');
|
||||
c2.width = RW * ZOOM; c2.height = RH * ZOOM;
|
||||
g2.fillStyle = '#0d0f16'; g2.fillRect(0, 0, c2.width, c2.height);
|
||||
g2.strokeStyle = '#2a2f3e'; g2.lineWidth = 1;
|
||||
g2.beginPath();
|
||||
g2.moveTo(c2.width / 2, 0); g2.lineTo(c2.width / 2, c2.height);
|
||||
g2.moveTo(0, c2.height / 2); g2.lineTo(c2.width, c2.height / 2);
|
||||
g2.stroke();
|
||||
const pl = (pts) => pts.map((p) => { const q = state.xform(p); return { x: q.x * ZOOM, y: q.y * ZOOM }; });
|
||||
strokePts(g2, pl(state.stab.oval[f]), '#3b4a63');
|
||||
strokePts(g2, pl(state.stab.outer[f]), '#4ade80');
|
||||
strokePts(g2, pl(state.stab.inner[f]), '#f87171');
|
||||
const ko = activeKey(state.keysOuter.keys, f);
|
||||
strokePts(g2, state.shapesOuter[ko.src].map((p) => ({ x: p.x * ZOOM, y: p.y * ZOOM })), '#fbbf24', 2);
|
||||
|
||||
// pane 3: the flat indexed render
|
||||
blit(el('cv-render'), renderFrame(f), ZOOM);
|
||||
}
|
||||
|
||||
function strokeRing(g, pts, w, h, color) {
|
||||
strokePts(g, pts.map((p) => ({ x: p.x * w, y: p.y * h })), color);
|
||||
}
|
||||
|
||||
function strokePts(g, pts, color, lw = 1) {
|
||||
g.strokeStyle = color; g.lineWidth = lw;
|
||||
g.beginPath();
|
||||
pts.forEach((p, i) => (i ? g.lineTo(p.x, p.y) : g.moveTo(p.x, p.y)));
|
||||
g.closePath(); g.stroke();
|
||||
}
|
||||
|
||||
function drawTimeline() {
|
||||
const c = el('cv-timeline'), g = c.getContext('2d');
|
||||
const N = state.dense.length;
|
||||
const W = Math.max(640, N * 8), H = 90;
|
||||
c.width = W; c.height = H;
|
||||
const px = W / N;
|
||||
g.fillStyle = '#0d0f16'; g.fillRect(0, 0, W, H);
|
||||
|
||||
const vel = state.keysOuter.velocity;
|
||||
const vmax = Math.max(...vel) || 1;
|
||||
g.strokeStyle = '#3b4a63'; g.lineWidth = 1; g.beginPath();
|
||||
vel.forEach((v, i) => {
|
||||
const x = i * px + px / 2, y = H - 18 - (v / vmax) * (H - 34);
|
||||
i ? g.lineTo(x, y) : g.moveTo(x, y);
|
||||
});
|
||||
g.stroke();
|
||||
|
||||
g.fillStyle = '#5b6478';
|
||||
for (const t of state.keysOuter.candidates) g.fillRect(t * px + px / 2 - 0.5, H - 18, 1, 6);
|
||||
|
||||
for (const k of state.keysOuter.keys) {
|
||||
g.fillStyle = '#fbbf24';
|
||||
g.fillRect(k.f * px + px / 2 - 1, 6, 2, H - 24); // f: what renders
|
||||
if (k.src !== k.f) {
|
||||
g.fillStyle = '#f87171';
|
||||
g.fillRect(k.src * px + px / 2 - 0.5, 6, 1, 10); // src: the true extreme
|
||||
}
|
||||
}
|
||||
g.fillStyle = '#e8eaf0';
|
||||
g.fillRect(state.frame * px + px / 2 - 0.5, 0, 1, H);
|
||||
}
|
||||
|
||||
function drawContactSheet() {
|
||||
const host = el('sheet');
|
||||
host.innerHTML = '';
|
||||
state.keysOuter.keys.forEach((k, i) => {
|
||||
const cell = document.createElement('div');
|
||||
cell.className = 'cell';
|
||||
const cv = document.createElement('canvas');
|
||||
blit(cv, renderFrame(k.f), 1);
|
||||
cv.style.width = '104px';
|
||||
const cap = document.createElement('span');
|
||||
cap.textContent = `f${k.f}${k.src !== k.f ? ` ←${k.src}` : ''}${state.hidden[i] ? ' ·closed' : ''}`;
|
||||
cell.append(cv, cap);
|
||||
cell.onclick = () => { state.frame = k.f; draw(); drawTimeline(); };
|
||||
host.append(cell);
|
||||
});
|
||||
}
|
||||
|
||||
/* ---------- export ---------- */
|
||||
|
||||
function exportTake() {
|
||||
const keys = state.keysOuter.keys;
|
||||
const take = {
|
||||
name: el('takename').value || 'line_01',
|
||||
frames: state.dense.length, width: RW, height: RH,
|
||||
exposure: opts().exposure,
|
||||
palette: PALETTE,
|
||||
slot: { x: RW / 2, y: RH / 2 },
|
||||
parts: [
|
||||
{ name: 'head', kind: 'plate', z: 0, interp: 'hold', keys: [{ f: 0, plate: 0 }] },
|
||||
{ name: 'mouth', kind: 'poly', z: 30, color: 'skin_dark', interp: 'hold',
|
||||
keys: keys.map((k) => ({ f: k.f, src: k.src, pts: state.shapesOuter[k.src] })) },
|
||||
{ name: 'mouth_in', kind: 'poly', z: 31, color: 'mouth_dark', interp: 'hold', parent: 'mouth',
|
||||
keys: keys.map((k, i) => state.hidden[i]
|
||||
? { f: k.f, hidden: true }
|
||||
: { f: k.f, src: k.src, pts: state.shapesInner[k.src] }) },
|
||||
],
|
||||
};
|
||||
const text = writeTake(take);
|
||||
const a = document.createElement('a');
|
||||
a.href = URL.createObjectURL(new Blob([text], { type: 'text/plain' }));
|
||||
a.download = `${take.name}.take`;
|
||||
a.click();
|
||||
status(`exported ${take.name}.take — ${keys.length} keys`, 'ok');
|
||||
}
|
||||
|
||||
/* ---------- wiring ---------- */
|
||||
|
||||
async function runFrames() {
|
||||
try {
|
||||
status('loading frames…');
|
||||
const images = await loadFrameSequence();
|
||||
if (!images.length) {
|
||||
status(`no frames found in ${el('framedir').value}/ — run extract.sh first`, 'err');
|
||||
return;
|
||||
}
|
||||
const { dense, missing } = await detectAll(images);
|
||||
state.images = images; state.dense = dense; state.source = 'video';
|
||||
el('scrub').max = dense.length - 1;
|
||||
state.frame = 0;
|
||||
rebuild();
|
||||
status(missing.length
|
||||
? `${images.length} frames · no face on ${missing.length} (held previous)`
|
||||
: `${images.length} frames detected`, missing.length ? 'warn' : 'ok');
|
||||
} catch (e) {
|
||||
status(`${e.message}`, 'err');
|
||||
console.error(e);
|
||||
}
|
||||
}
|
||||
|
||||
function runSynthetic() {
|
||||
state.images = [];
|
||||
state.dense = synthDense(72);
|
||||
state.source = 'synthetic';
|
||||
el('scrub').max = 71;
|
||||
state.frame = 0;
|
||||
rebuild();
|
||||
status('synthetic take — no video needed; exercises the whole chain below detection', 'ok');
|
||||
}
|
||||
|
||||
for (const id of ['verts','minHold','distThresh','velSmooth','smoothWin','exposure','apertureThresh']) {
|
||||
el(id).addEventListener('input', () => {
|
||||
el(id + 'v').textContent = id === 'distThresh' ? (+el(id).value / 10).toFixed(1)
|
||||
: id === 'apertureThresh' ? (+el(id).value / 1000).toFixed(3)
|
||||
: el(id).value;
|
||||
rebuild();
|
||||
});
|
||||
el(id + 'v').textContent = el(id).value;
|
||||
}
|
||||
el('scrub').addEventListener('input', (e) => { state.frame = +e.target.value; draw(); drawTimeline(); });
|
||||
el('btn-frames').onclick = runFrames;
|
||||
el('btn-synth').onclick = runSynthetic;
|
||||
el('btn-export').onclick = exportTake;
|
||||
el('btn-play').onclick = () => {
|
||||
state.playing = !state.playing;
|
||||
el('btn-play').textContent = state.playing ? 'Stop' : 'Play';
|
||||
if (state.playing) tick();
|
||||
};
|
||||
|
||||
let last = 0;
|
||||
function tick(ts = 0) {
|
||||
if (!state.playing) return;
|
||||
if (ts - last > 1000 / 24) {
|
||||
last = ts;
|
||||
state.frame = (state.frame + 1) % state.dense.length;
|
||||
el('scrub').value = state.frame;
|
||||
draw(); drawTimeline();
|
||||
}
|
||||
requestAnimationFrame(tick);
|
||||
}
|
||||
|
||||
PALETTE.forEach((p) => {
|
||||
const sw = document.createElement('label');
|
||||
sw.className = 'sw';
|
||||
const inp = document.createElement('input');
|
||||
inp.type = 'color'; inp.value = p.hex;
|
||||
inp.oninput = () => { p.hex = inp.value; if (state.dense) { draw(); drawContactSheet(); } };
|
||||
sw.append(inp, document.createTextNode(p.name));
|
||||
el('palette').append(sw);
|
||||
});
|
||||
|
||||
status('ready — "Synthetic take" works with no video; "Load frames" reads ./frames/');
|
||||
54
js/landmarks.js
Normal file
54
js/landmarks.js
Normal file
|
|
@ -0,0 +1,54 @@
|
|||
// MediaPipe FaceLandmarker index tables.
|
||||
// Ring arrays are ORDERED traversals, not raw connection sets: vertex position
|
||||
// within a ring is the vertex's identity, and every downstream stage depends on
|
||||
// that ordering being stable. See docs/roto-puppet.md, "Fixed topology".
|
||||
|
||||
// Rigid landmarks for the similarity fit. Eye corners, nose bridge, nose tip.
|
||||
// Nothing here may be a feature that moves under performance: including the
|
||||
// mouth or brows bleeds performance into the stabilization.
|
||||
export const RIGID = [33, 133, 362, 263, 168, 6, 1];
|
||||
|
||||
// Outer lip ring, clockwise from the right corner over the top.
|
||||
// index 0 = right corner, 5 = top centre, 10 = left corner, 15 = bottom centre.
|
||||
export const LIPS_OUTER = [
|
||||
61, 185, 40, 39, 37, 0, 267, 269, 270, 409,
|
||||
291, 375, 321, 405, 314, 17, 84, 181, 91, 146,
|
||||
];
|
||||
|
||||
// Inner lip ring, same orientation and the same four cardinal positions.
|
||||
export const LIPS_INNER = [
|
||||
78, 191, 80, 81, 82, 13, 312, 311, 310, 415,
|
||||
308, 324, 318, 402, 317, 14, 87, 178, 88, 95,
|
||||
];
|
||||
|
||||
// Inner upper / lower lip centres. Their separation is the aperture signal that
|
||||
// decides whether the mouth interior is present at all.
|
||||
export const APERTURE = [13, 14];
|
||||
|
||||
// Face oval, used only to derive the placeholder plate in v1.
|
||||
export const FACE_OVAL = [
|
||||
10, 338, 297, 332, 284, 251, 389, 356, 454, 323, 361, 288,
|
||||
397, 365, 379, 378, 400, 377, 152, 148, 176, 149, 150, 136,
|
||||
172, 58, 132, 93, 234, 127, 162, 21, 54, 103, 67, 109,
|
||||
];
|
||||
|
||||
// Eye corners, for the calibration box and for reporting fit residual.
|
||||
export const EYE_INNER = [133, 362];
|
||||
|
||||
// Pick `n` slots from a ring of `len` by even spacing. Returns RING POSITIONS,
|
||||
// not landmark ids: positions are the vertex identity downstream, and mapping ids
|
||||
// back to positions with indexOf would silently pick the wrong slot if a table
|
||||
// ever repeated an id.
|
||||
//
|
||||
// For even n this naturally lands on the cardinal positions (corners and lip
|
||||
// centres) of a 20-point ring. Fixed indices, never adaptive decimation: the
|
||||
// vertex at slot k means the same thing on every frame of the shot.
|
||||
export function subsampleSlots(len, n) {
|
||||
const out = [];
|
||||
for (let k = 0; k < n; k++) out.push(Math.round((k * len) / n) % len);
|
||||
return out;
|
||||
}
|
||||
|
||||
export function subsampleRing(ring, n) {
|
||||
return subsampleSlots(ring.length, n).map((s) => ring[s]);
|
||||
}
|
||||
107
js/mathutil.js
Normal file
107
js/mathutil.js
Normal file
|
|
@ -0,0 +1,107 @@
|
|||
// 2D similarity transforms and temporal smoothing.
|
||||
|
||||
// Least-squares similarity (translation + rotation + uniform scale, 4 DOF)
|
||||
// mapping P onto Q. Closed form; no iteration.
|
||||
//
|
||||
// Deliberately NOT affine or homography: the extra degrees of freedom absorb
|
||||
// out-of-plane head rotation as shear/perspective and smear it into the mouth.
|
||||
// Four DOF removes exactly translation, roll and depth-scale, and leaves yaw and
|
||||
// pitch as a measurable residual.
|
||||
export function fitSimilarity(P, Q) {
|
||||
const n = P.length;
|
||||
let pcx = 0, pcy = 0, qcx = 0, qcy = 0;
|
||||
for (let i = 0; i < n; i++) {
|
||||
pcx += P[i].x; pcy += P[i].y;
|
||||
qcx += Q[i].x; qcy += Q[i].y;
|
||||
}
|
||||
pcx /= n; pcy /= n; qcx /= n; qcy /= n;
|
||||
|
||||
let a = 0, b = 0, norm = 0;
|
||||
for (let i = 0; i < n; i++) {
|
||||
const px = P[i].x - pcx, py = P[i].y - pcy;
|
||||
const qx = Q[i].x - qcx, qy = Q[i].y - qcy;
|
||||
a += px * qx + py * qy; // dot
|
||||
b += px * qy - py * qx; // cross
|
||||
norm += px * px + py * py;
|
||||
}
|
||||
const theta = Math.atan2(b, a);
|
||||
const s = norm > 1e-12 ? Math.hypot(a, b) / norm : 1;
|
||||
const c = Math.cos(theta), sn = Math.sin(theta);
|
||||
return {
|
||||
s, theta,
|
||||
tx: qcx - s * (c * pcx - sn * pcy),
|
||||
ty: qcy - s * (sn * pcx + c * pcy),
|
||||
};
|
||||
}
|
||||
|
||||
export function applySim(tf, p) {
|
||||
const c = Math.cos(tf.theta), sn = Math.sin(tf.theta);
|
||||
return {
|
||||
x: tf.s * (c * p.x - sn * p.y) + tf.tx,
|
||||
y: tf.s * (sn * p.x + c * p.y) + tf.ty,
|
||||
};
|
||||
}
|
||||
|
||||
export function applySimAll(tf, pts) {
|
||||
return pts.map((p) => applySim(tf, p));
|
||||
}
|
||||
|
||||
// Residual RMS after the fit, in the units of Q. Rises with out-of-plane
|
||||
// rotation, so it is the signal for "this section is not stabilisable".
|
||||
export function fitResidual(tf, P, Q) {
|
||||
let acc = 0;
|
||||
for (let i = 0; i < P.length; i++) {
|
||||
const m = applySim(tf, P[i]);
|
||||
acc += (m.x - Q[i].x) ** 2 + (m.y - Q[i].y) ** 2;
|
||||
}
|
||||
return Math.sqrt(acc / P.length);
|
||||
}
|
||||
|
||||
// Generalised Procrustes: the reference is the MEAN rigid configuration over the
|
||||
// shot, not frame zero, so no single frame's idiosyncrasies get baked into every
|
||||
// other frame. Three passes is plenty.
|
||||
export function procrustesMean(framesRigid, iters = 3) {
|
||||
let ref = framesRigid[0].map((p) => ({ x: p.x, y: p.y }));
|
||||
for (let it = 0; it < iters; it++) {
|
||||
const acc = ref.map(() => ({ x: 0, y: 0 }));
|
||||
for (const rig of framesRigid) {
|
||||
const tf = fitSimilarity(rig, ref);
|
||||
const moved = applySimAll(tf, rig);
|
||||
for (let i = 0; i < acc.length; i++) { acc[i].x += moved[i].x; acc[i].y += moved[i].y; }
|
||||
}
|
||||
ref = acc.map((p) => ({ x: p.x / framesRigid.length, y: p.y / framesRigid.length }));
|
||||
}
|
||||
return ref;
|
||||
}
|
||||
|
||||
function movingAverage(vals, win) {
|
||||
if (win <= 1) return vals.slice();
|
||||
const half = Math.floor(win / 2), out = new Array(vals.length);
|
||||
for (let i = 0; i < vals.length; i++) {
|
||||
let acc = 0, cnt = 0;
|
||||
for (let j = i - half; j <= i + half; j++) {
|
||||
const k = Math.min(vals.length - 1, Math.max(0, j));
|
||||
acc += vals[k]; cnt++;
|
||||
}
|
||||
out[i] = acc / cnt;
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
export { movingAverage };
|
||||
|
||||
// Smooth the four transform parameters, NEVER the contour. Landmark jitter of a
|
||||
// pixel is smeared into the mouth by the inverse transform, so the transform is
|
||||
// where the low-pass belongs; smoothing the contour would destroy the
|
||||
// performance, which is the entire asset.
|
||||
// Angles are smoothed as (cos, sin) so wrapping cannot produce a spike.
|
||||
export function smoothTransforms(tfs, win) {
|
||||
const c = movingAverage(tfs.map((t) => Math.cos(t.theta)), win);
|
||||
const sn = movingAverage(tfs.map((t) => Math.sin(t.theta)), win);
|
||||
const s = movingAverage(tfs.map((t) => t.s), win);
|
||||
const tx = movingAverage(tfs.map((t) => t.tx), win);
|
||||
const ty = movingAverage(tfs.map((t) => t.ty), win);
|
||||
return tfs.map((_, i) => ({
|
||||
theta: Math.atan2(sn[i], c[i]), s: s[i], tx: tx[i], ty: ty[i],
|
||||
}));
|
||||
}
|
||||
108
js/pipeline.js
Normal file
108
js/pipeline.js
Normal file
|
|
@ -0,0 +1,108 @@
|
|||
// Analysis: dense track -> stabilised head-local contours -> selected keys.
|
||||
// All policy lives here, never in the renderer. See docs/roto-puppet.md,
|
||||
// "The take is the contract".
|
||||
|
||||
import { RIGID, LIPS_OUTER, LIPS_INNER, APERTURE, FACE_OVAL, EYE_INNER, subsampleSlots } from './landmarks.js';
|
||||
import { fitSimilarity, applySimAll, applySim, fitResidual, procrustesMean, smoothTransforms, movingAverage } from './mathutil.js';
|
||||
|
||||
const pick = (lm, idx) => idx.map((i) => ({ x: lm[i].x, y: lm[i].y }));
|
||||
|
||||
// Stage 1-3: fit the rigid transform per frame, smooth its parameters, then map
|
||||
// every contour through it into the reference frame. The result is head-local:
|
||||
// translation, roll and depth-scale of the head are gone.
|
||||
export function stabilize(dense, smoothWin) {
|
||||
const rigid = dense.map((f) => pick(f, RIGID));
|
||||
const ref = procrustesMean(rigid);
|
||||
const raw = rigid.map((r) => fitSimilarity(r, ref));
|
||||
const tfs = smoothTransforms(raw, smoothWin);
|
||||
|
||||
return {
|
||||
ref,
|
||||
transforms: tfs,
|
||||
// Residual rises with out-of-plane rotation, which no 2D similarity can
|
||||
// remove. High values mean this section wants a different head plate.
|
||||
residual: tfs.map((tf, i) => fitResidual(tf, rigid[i], ref)),
|
||||
outer: dense.map((f, i) => applySimAll(tfs[i], pick(f, LIPS_OUTER))),
|
||||
inner: dense.map((f, i) => applySimAll(tfs[i], pick(f, LIPS_INNER))),
|
||||
oval: dense.map((f, i) => applySimAll(tfs[i], pick(f, FACE_OVAL))),
|
||||
eyes: dense.map((f, i) => applySimAll(tfs[i], pick(f, EYE_INNER))),
|
||||
aperture: dense.map((f, i) => {
|
||||
const a = applySimAll(tfs[i], pick(f, APERTURE));
|
||||
return Math.hypot(a[0].x - a[1].x, a[0].y - a[1].y);
|
||||
}),
|
||||
};
|
||||
}
|
||||
|
||||
// Stage 4: fixed-index subsample of a stabilised ring, then map from normalised
|
||||
// face space into character raster space.
|
||||
export function toRasterRing(stabRing, ringTable, n, xform) {
|
||||
return subsampleSlots(ringTable.length, n).map((s) => xform(stabRing[s]));
|
||||
}
|
||||
|
||||
// Stage 6: key selection.
|
||||
//
|
||||
// Keys go on velocity MINIMA, not on distance thresholds. A threshold fires at
|
||||
// the frame it was crossed - partway through a transition - so every pose lands
|
||||
// mushy and late. A minimum is where the shape is momentarily parked, which is
|
||||
// the pose a viewer actually reads.
|
||||
//
|
||||
// Minima alone are not enough: during a long hold the velocity wobbles near zero
|
||||
// and produces a key per wobble. So a candidate minimum is only accepted if the
|
||||
// shape has actually moved since the last accepted key (distThresh) and the
|
||||
// minimum hold has elapsed (minHold).
|
||||
export function selectKeys(shapes, opts) {
|
||||
const { minHold, distThresh, velSmooth, exposure } = opts;
|
||||
const N = shapes.length;
|
||||
if (N === 0) return { keys: [], velocity: [], candidates: [] };
|
||||
|
||||
const vel = new Array(N).fill(0);
|
||||
for (let t = 1; t < N; t++) {
|
||||
let acc = 0;
|
||||
for (let i = 0; i < shapes[t].length; i++) {
|
||||
acc += Math.hypot(shapes[t][i].x - shapes[t - 1][i].x, shapes[t][i].y - shapes[t - 1][i].y);
|
||||
}
|
||||
vel[t] = acc / shapes[t].length;
|
||||
}
|
||||
const sv = movingAverage(vel, velSmooth);
|
||||
|
||||
const candidates = [];
|
||||
for (let t = 1; t < N - 1; t++) {
|
||||
if (sv[t] <= sv[t - 1] && sv[t] <= sv[t + 1]) candidates.push(t);
|
||||
}
|
||||
|
||||
const shapeDist = (a, b) => {
|
||||
let acc = 0;
|
||||
for (let i = 0; i < a.length; i++) acc += Math.hypot(a[i].x - b[i].x, a[i].y - b[i].y);
|
||||
return acc / a.length;
|
||||
};
|
||||
|
||||
const accepted = [0];
|
||||
for (const t of candidates) {
|
||||
const last = accepted[accepted.length - 1];
|
||||
if (t - last < minHold) continue;
|
||||
if (shapeDist(shapes[t], shapes[last]) < distThresh) continue;
|
||||
accepted.push(t);
|
||||
}
|
||||
|
||||
// Snap onto the exposure grid. f is what renders; src is provenance.
|
||||
const keys = [];
|
||||
for (const src of accepted) {
|
||||
const f = Math.round(src / exposure) * exposure;
|
||||
const prev = keys[keys.length - 1];
|
||||
if (prev && prev.f === f) {
|
||||
// Two extremes collapsed onto one grid slot: keep the stronger one.
|
||||
if (sv[src] < sv[prev.src]) { prev.src = src; prev.frame = src; }
|
||||
continue;
|
||||
}
|
||||
keys.push({ f, src, frame: src });
|
||||
}
|
||||
return { keys, velocity: sv, candidates };
|
||||
}
|
||||
|
||||
// Resolve which key is live on a given output frame under interp=hold.
|
||||
// "Most recent key at or before f" - lookup, not policy.
|
||||
export function activeKey(keys, f) {
|
||||
let hit = keys[0];
|
||||
for (const k of keys) { if (k.f <= f) hit = k; else break; }
|
||||
return hit;
|
||||
}
|
||||
83
js/raster.js
Normal file
83
js/raster.js
Normal file
|
|
@ -0,0 +1,83 @@
|
|||
// Indexed flat-fill rasteriser.
|
||||
//
|
||||
// Canvas2D antialiases path fills, and antialiasing is exactly what the target
|
||||
// idiom does not have: Animator Pro fills polygons into a 256-colour indexed
|
||||
// raster with hard edges (csd_render_poly). A preview that antialiases would
|
||||
// misrepresent the look it exists to judge, so this writes palette indices into
|
||||
// a byte buffer with an even-odd scanline fill and expands to RGBA only at the
|
||||
// very end.
|
||||
|
||||
export class IndexedRaster {
|
||||
constructor(w, h) {
|
||||
this.w = w; this.h = h;
|
||||
this.buf = new Uint8Array(w * h);
|
||||
}
|
||||
|
||||
clear(index) { this.buf.fill(index); }
|
||||
|
||||
// Even-odd scanline fill. Samples at pixel centres (y + 0.5), so a polygon
|
||||
// edge landing exactly on a pixel boundary resolves consistently.
|
||||
fillPoly(pts, index) {
|
||||
const n = pts.length;
|
||||
if (n < 3) return;
|
||||
let minY = Infinity, maxY = -Infinity;
|
||||
for (const p of pts) { if (p.y < minY) minY = p.y; if (p.y > maxY) maxY = p.y; }
|
||||
const y0 = Math.max(0, Math.ceil(minY - 0.5));
|
||||
const y1 = Math.min(this.h - 1, Math.floor(maxY - 0.5) + 1);
|
||||
const xs = [];
|
||||
for (let y = y0; y <= y1; y++) {
|
||||
const sy = y + 0.5;
|
||||
xs.length = 0;
|
||||
for (let i = 0; i < n; i++) {
|
||||
const a = pts[i], b = pts[(i + 1) % n];
|
||||
if (a.y === b.y) continue;
|
||||
const lo = Math.min(a.y, b.y), hi = Math.max(a.y, b.y);
|
||||
if (sy < lo || sy >= hi) continue;
|
||||
xs.push(a.x + ((sy - a.y) / (b.y - a.y)) * (b.x - a.x));
|
||||
}
|
||||
if (xs.length < 2) continue;
|
||||
xs.sort((p, q) => p - q);
|
||||
for (let k = 0; k + 1 < xs.length; k += 2) {
|
||||
const xa = Math.max(0, Math.ceil(xs[k] - 0.5));
|
||||
const xb = Math.min(this.w - 1, Math.floor(xs[k + 1] - 0.5));
|
||||
const row = y * this.w;
|
||||
for (let x = xa; x <= xb; x++) this.buf[row + x] = index;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
fillDisc(cx, cy, r, index) {
|
||||
const rr = r * r;
|
||||
const y0 = Math.max(0, Math.floor(cy - r)), y1 = Math.min(this.h - 1, Math.ceil(cy + r));
|
||||
const x0 = Math.max(0, Math.floor(cx - r)), x1 = Math.min(this.w - 1, Math.ceil(cx + r));
|
||||
for (let y = y0; y <= y1; y++) {
|
||||
for (let x = x0; x <= x1; x++) {
|
||||
const dx = x + 0.5 - cx, dy = y + 0.5 - cy;
|
||||
if (dx * dx + dy * dy <= rr) this.buf[y * this.w + x] = index;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Expand indices through the palette into an ImageData at integer zoom.
|
||||
// Nearest-neighbour by construction, so no filtering softens the result.
|
||||
toImageData(palette, zoom = 1) {
|
||||
const W = this.w * zoom, H = this.h * zoom;
|
||||
const img = new ImageData(W, H);
|
||||
const d = img.data;
|
||||
const rgb = palette.map(hexToRgb);
|
||||
for (let y = 0; y < H; y++) {
|
||||
const srow = Math.floor(y / zoom) * this.w;
|
||||
for (let x = 0; x < W; x++) {
|
||||
const c = rgb[this.buf[srow + Math.floor(x / zoom)]] || [255, 0, 255];
|
||||
const o = (y * W + x) * 4;
|
||||
d[o] = c[0]; d[o + 1] = c[1]; d[o + 2] = c[2]; d[o + 3] = 255;
|
||||
}
|
||||
}
|
||||
return img;
|
||||
}
|
||||
}
|
||||
|
||||
export function hexToRgb(hex) {
|
||||
const s = hex.replace('#', '');
|
||||
return [parseInt(s.slice(0, 2), 16), parseInt(s.slice(2, 4), 16), parseInt(s.slice(4, 6), 16)];
|
||||
}
|
||||
163
js/selftest.js
Normal file
163
js/selftest.js
Normal file
|
|
@ -0,0 +1,163 @@
|
|||
// Assertions over the stages below detection. Runs in the browser so the exact
|
||||
// module graph the tool uses is what gets tested.
|
||||
//
|
||||
// The ring-simplicity check exists because "fixed topology" is load-bearing in
|
||||
// docs/roto-puppet.md: because hold parts CUT between poses rather than
|
||||
// interpolating, a ring whose vertex order is wrong self-intersects and renders
|
||||
// as blocks meeting at corners. It is invisible at some vertex counts and obvious
|
||||
// at others, so it needs an assertion rather than an eyeball.
|
||||
|
||||
import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, subsampleSlots, subsampleRing } from './landmarks.js';
|
||||
import { fitSimilarity, applySim, procrustesMean, smoothTransforms } from './mathutil.js';
|
||||
import { stabilize, toRasterRing, selectKeys, activeKey } from './pipeline.js';
|
||||
import { IndexedRaster, hexToRgb } from './raster.js';
|
||||
import { writeTake } from './take.js';
|
||||
import { synthDense } from './synth.js';
|
||||
|
||||
const results = [];
|
||||
const ok = (name, cond, detail = '') => results.push({ name, pass: !!cond, detail });
|
||||
|
||||
/* ---- geometry helpers ---- */
|
||||
|
||||
function segmentsCross(a, b, c, d) {
|
||||
const o = (p, q, r) => Math.sign((q.x - p.x) * (r.y - p.y) - (q.y - p.y) * (r.x - p.x));
|
||||
const o1 = o(a, b, c), o2 = o(a, b, d), o3 = o(c, d, a), o4 = o(c, d, b);
|
||||
return o1 !== o2 && o3 !== o4 && o1 !== 0 && o2 !== 0 && o3 !== 0 && o4 !== 0;
|
||||
}
|
||||
|
||||
// A closed ring is simple if no pair of non-adjacent edges crosses.
|
||||
function ringSelfIntersections(pts) {
|
||||
const n = pts.length, hits = [];
|
||||
for (let i = 0; i < n; i++) {
|
||||
for (let j = i + 1; j < n; j++) {
|
||||
if (i === j || (j + 1) % n === i || (i + 1) % n === j) continue;
|
||||
if (segmentsCross(pts[i], pts[(i + 1) % n], pts[j], pts[(j + 1) % n])) hits.push([i, j]);
|
||||
}
|
||||
}
|
||||
return hits;
|
||||
}
|
||||
|
||||
const spreadX = (frames, slot) => {
|
||||
const xs = frames.map((f) => f[slot].x);
|
||||
return Math.max(...xs) - Math.min(...xs);
|
||||
};
|
||||
|
||||
/* ---- the tests ---- */
|
||||
|
||||
export function run() {
|
||||
results.length = 0;
|
||||
|
||||
// tables
|
||||
ok('LIPS_OUTER has 20 distinct ids', new Set(LIPS_OUTER).size === 20);
|
||||
ok('LIPS_INNER has 20 distinct ids', new Set(LIPS_INNER).size === 20);
|
||||
ok('FACE_OVAL has 36 distinct ids', new Set(FACE_OVAL).size === 36);
|
||||
ok('RIGID excludes every lip vertex',
|
||||
!RIGID.some((i) => LIPS_OUTER.includes(i) || LIPS_INNER.includes(i)),
|
||||
'a moving feature in the rigid set bleeds performance into stabilisation');
|
||||
|
||||
// subsampling preserves order and count at every budget
|
||||
for (let n = 4; n <= 16; n += 2) {
|
||||
const s = subsampleSlots(20, n);
|
||||
const mono = s.every((v, i) => i === 0 || v > s[i - 1]);
|
||||
ok(`subsampleSlots(20,${n}) is strictly increasing, n=${n}`, mono && s.length === n, s.join(','));
|
||||
}
|
||||
ok('subsampleRing agrees with subsampleSlots',
|
||||
subsampleRing(LIPS_OUTER, 8).join(',') === subsampleSlots(20, 8).map((s) => LIPS_OUTER[s]).join(','));
|
||||
|
||||
const dense = synthDense(72);
|
||||
|
||||
// rings must be simple at EVERY vertex budget, on every frame
|
||||
for (const [label, table] of [['outer', LIPS_OUTER], ['inner', LIPS_INNER]]) {
|
||||
let worst = null;
|
||||
for (let n = 4; n <= 16 && !worst; n += 2) {
|
||||
const slots = subsampleSlots(table.length, n);
|
||||
for (let f = 0; f < dense.length; f++) {
|
||||
const pts = slots.map((s) => dense[f][table[s]]);
|
||||
const hits = ringSelfIntersections(pts);
|
||||
if (hits.length) { worst = `verts=${n} frame=${f} edges ${JSON.stringify(hits[0])}`; break; }
|
||||
}
|
||||
}
|
||||
ok(`${label} ring is simple at every vertex budget`, !worst, worst || '');
|
||||
}
|
||||
|
||||
// similarity fit recovers a known transform
|
||||
const src = [{ x: 0, y: 0 }, { x: 1, y: 0 }, { x: 0, y: 1 }, { x: 2, y: 3 }];
|
||||
const truth = { s: 1.7, theta: 0.6, tx: 4, ty: -2 };
|
||||
const dst = src.map((p) => applySim(truth, p));
|
||||
const got = fitSimilarity(src, dst);
|
||||
ok('fitSimilarity recovers a known transform',
|
||||
Math.abs(got.s - truth.s) < 1e-9 && Math.abs(got.theta - truth.theta) < 1e-9 &&
|
||||
Math.abs(got.tx - truth.tx) < 1e-9 && Math.abs(got.ty - truth.ty) < 1e-9,
|
||||
`s=${got.s.toFixed(6)} th=${got.theta.toFixed(6)}`);
|
||||
|
||||
// stabilisation: head motion out, mouth motion kept
|
||||
const stab = stabilize(dense, 1);
|
||||
const rawSpread = spreadX(dense, 133);
|
||||
const stabSpread = (() => {
|
||||
const xs = stab.eyes.map((e) => e[0].x);
|
||||
return Math.max(...xs) - Math.min(...xs);
|
||||
})();
|
||||
ok('stabilisation removes >90% of head translation',
|
||||
stabSpread < rawSpread * 0.1, `raw ${rawSpread.toFixed(4)} -> ${stabSpread.toFixed(4)}`);
|
||||
const apRange = Math.max(...stab.aperture) - Math.min(...stab.aperture);
|
||||
ok('stabilisation preserves mouth motion', apRange > 0.05, `aperture range ${apRange.toFixed(4)}`);
|
||||
|
||||
// key selection
|
||||
const xf = (p) => ({ x: p.x * 320, y: p.y * 200 });
|
||||
const shapes = stab.outer.map((r) => toRasterRing(r, LIPS_OUTER, 8, xf));
|
||||
const sel = selectKeys(shapes, { minHold: 2, distThresh: 0.6, velSmooth: 3, exposure: 2 });
|
||||
ok('keys are strictly increasing in f', sel.keys.every((k, i) => i === 0 || k.f > sel.keys[i - 1].f));
|
||||
ok('keys respect the minimum hold',
|
||||
sel.keys.every((k, i) => i === 0 || k.src - sel.keys[i - 1].src >= 2));
|
||||
ok('keys land on the exposure grid', sel.keys.every((k) => k.f % 2 === 0));
|
||||
ok('selection reduces candidates', sel.keys.length < sel.candidates.length,
|
||||
`${sel.candidates.length} candidates -> ${sel.keys.length} keys`);
|
||||
ok('first key is frame 0', sel.keys[0].f === 0);
|
||||
ok('activeKey holds between keys',
|
||||
activeKey(sel.keys, sel.keys[1].f - 1).f === sel.keys[0].f);
|
||||
|
||||
// rasteriser: indexed, hard-edged, no blending
|
||||
const r = new IndexedRaster(64, 48);
|
||||
r.clear(0);
|
||||
r.fillPoly([{ x: 8, y: 8 }, { x: 56, y: 8 }, { x: 56, y: 40 }, { x: 8, y: 40 }], 2);
|
||||
const present = new Set(r.buf);
|
||||
ok('raster contains only written indices', present.size === 2 && present.has(0) && present.has(2),
|
||||
`indices ${[...present].join(',')}`);
|
||||
let count = 0;
|
||||
for (const v of r.buf) if (v === 2) count++;
|
||||
ok('axis-aligned rect fills the exact pixel count', count === 48 * 32, `${count} vs ${48 * 32}`);
|
||||
const pal = ['#000000', '#ffffff', '#ff8800'];
|
||||
const img = r.toImageData(pal, 2);
|
||||
const seen = new Set();
|
||||
for (let i = 0; i < img.data.length; i += 4) {
|
||||
seen.add(`${img.data[i]},${img.data[i + 1]},${img.data[i + 2]}`);
|
||||
}
|
||||
const allowed = new Set(pal.map((h) => hexToRgb(h).join(',')));
|
||||
ok('palette expansion introduces no intermediate colours',
|
||||
[...seen].every((c) => allowed.has(c)), `${seen.size} distinct colours`);
|
||||
|
||||
// take writer round-trip
|
||||
const take = {
|
||||
name: 'test', frames: 72, width: 320, height: 200, exposure: 2,
|
||||
palette: [{ name: 'bg' }, { name: 'skin' }],
|
||||
slot: { x: 160, y: 100 },
|
||||
parts: [
|
||||
{ name: 'head', kind: 'plate', z: 0, interp: 'hold', keys: [{ f: 0, plate: 0 }] },
|
||||
{ name: 'mouth', kind: 'poly', z: 30, color: 'skin', interp: 'hold',
|
||||
keys: sel.keys.map((k) => ({ f: k.f, src: k.src, pts: shapes[k.src] })) },
|
||||
],
|
||||
};
|
||||
const text = writeTake(take);
|
||||
const keyLines = text.split('\n').filter((l) => l.startsWith('key') && l.includes('n='));
|
||||
ok('every key line declares n= matching its point count',
|
||||
keyLines.every((l) => {
|
||||
const n = +l.match(/n=(\d+)/)[1];
|
||||
const pts = l.split(/n=\d+\s+/)[1].trim().split(/\s+/);
|
||||
return pts.length === n;
|
||||
}), `${keyLines.length} key lines`);
|
||||
ok('take declares a plate and a part table',
|
||||
/^plate\s+0/m.test(text) && /^part\s+mouth/m.test(text));
|
||||
ok('coordinates are integers', !/-?\d+\.\d/.test(text.split('\n').filter((l) => l.startsWith('key')).join('')));
|
||||
|
||||
return results;
|
||||
}
|
||||
69
js/synth.js
Normal file
69
js/synth.js
Normal file
|
|
@ -0,0 +1,69 @@
|
|||
// Synthetic landmark frames, shaped exactly like FaceLandmarker output.
|
||||
//
|
||||
// Exists so the whole chain downstream of detection - Procrustes, smoothing,
|
||||
// stabilisation, key selection, rasterising, take writing - can be exercised and
|
||||
// verified without a video file. A synthetic face is also the only way to test
|
||||
// stabilisation against a KNOWN head motion, since real footage gives no ground
|
||||
// truth to compare against.
|
||||
import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, EYE_INNER } from './landmarks.js';
|
||||
|
||||
const NUM = 478;
|
||||
|
||||
export function synthDense(nFrames = 72) {
|
||||
const frames = [];
|
||||
for (let t = 0; t < nFrames; t++) {
|
||||
const pts = new Array(NUM);
|
||||
for (let i = 0; i < NUM; i++) pts[i] = { x: 0.5, y: 0.5, z: 0 };
|
||||
|
||||
// Known head motion: drift, sway, roll and a slow scale change, plus a
|
||||
// little per-frame jitter so transform smoothing has something to remove.
|
||||
const ph = t / nFrames;
|
||||
const hx = 0.5 + 0.045 * Math.sin(ph * Math.PI * 2) + (Math.random() - 0.5) * 0.002;
|
||||
const hy = 0.5 + 0.02 * Math.cos(ph * Math.PI * 3) + (Math.random() - 0.5) * 0.002;
|
||||
const roll = 0.18 * Math.sin(ph * Math.PI * 2.5);
|
||||
const scale = 1 + 0.06 * Math.sin(ph * Math.PI * 1.5);
|
||||
const cr = Math.cos(roll), sr = Math.sin(roll);
|
||||
const place = (i, lx, ly) => {
|
||||
const sx = lx * scale, sy = ly * scale;
|
||||
pts[i] = { x: hx + cr * sx - sr * sy, y: hy + sr * sx + cr * sy, z: 0 };
|
||||
};
|
||||
|
||||
// Mouth opens in four sustained beats with holds between, so key selection
|
||||
// has genuine extremes and genuine plateaux to find.
|
||||
const beat = Math.floor(t / 9) % 4;
|
||||
const target = [0.004, 0.05, 0.022, 0.0];
|
||||
const openAmt = target[beat];
|
||||
const wide = 0.10 + (beat === 1 ? 0.012 : beat === 3 ? -0.008 : 0);
|
||||
|
||||
place(RIGID[0], -0.075, -0.045); place(RIGID[1], -0.028, -0.043);
|
||||
place(RIGID[2], 0.028, -0.043); place(RIGID[3], 0.075, -0.045);
|
||||
place(RIGID[4], 0.000, -0.050); place(RIGID[5], 0.000, -0.020);
|
||||
place(RIGID[6], 0.000, 0.012);
|
||||
place(EYE_INNER[0], -0.028, -0.043); place(EYE_INNER[1], 0.028, -0.043);
|
||||
|
||||
// Lip rings as ellipse arcs, traversed so ring ORDER matches the tables:
|
||||
// slot 0 = right corner, 5 = top centre, 10 = left corner, 15 = bottom
|
||||
// centre, with y growing downward. Getting this convention wrong swaps two
|
||||
// opposite vertices and the ring self-intersects into a bowtie - see the
|
||||
// ring-simplicity assertion in selftest.
|
||||
const ring = (table, rx, ry, cy) => {
|
||||
const n = table.length;
|
||||
for (let k = 0; k < n; k++) {
|
||||
const a = -(k / n) * Math.PI * 2;
|
||||
place(table[k], rx * Math.cos(a), cy + ry * Math.sin(a));
|
||||
}
|
||||
};
|
||||
ring(LIPS_OUTER, wide / 2, 0.012 + openAmt * 0.6, 0.075);
|
||||
// APERTURE (13, 14) are slots 5 and 15 of the inner ring, so the ring itself
|
||||
// places them at the vertical extremes. Writing them again afterwards is what
|
||||
// produced the bowtie; the aperture is simply the inner ring's height.
|
||||
ring(LIPS_INNER, wide / 2.6, 0.001 + openAmt, 0.075);
|
||||
|
||||
for (let k = 0; k < FACE_OVAL.length; k++) {
|
||||
const a = -Math.PI / 2 + (k / FACE_OVAL.length) * Math.PI * 2;
|
||||
place(FACE_OVAL[k], 0.105 * Math.cos(a), 0.145 * Math.sin(a) + 0.01);
|
||||
}
|
||||
frames.push(pts);
|
||||
}
|
||||
return frames;
|
||||
}
|
||||
41
js/take.js
Normal file
41
js/take.js
Normal file
|
|
@ -0,0 +1,41 @@
|
|||
// Take-file writer. Format is specified in docs/roto-puppet.md, "The take
|
||||
// format". Line-oriented text on purpose: Poco can parse it with fopen/fgets
|
||||
// from poco/src/safefile.c and strtok/atoi/atof from poco/src/strlib.c, so the
|
||||
// Animator Pro render script needs no new native code.
|
||||
|
||||
const r = (v) => Math.round(v);
|
||||
|
||||
export function writeTake(take) {
|
||||
const L = [];
|
||||
L.push(`take name=${take.name} frames=${take.frames} width=${take.width} height=${take.height} exposure=${take.exposure}`);
|
||||
L.push('');
|
||||
take.palette.forEach((p, i) => L.push(`pal ${i} ${p.name.padEnd(11)} ${i}`));
|
||||
L.push('');
|
||||
// v1 emits a single frozen plate derived from the face oval. A real project
|
||||
// replaces this with hand-drawn angles referenced by cel frame; the record
|
||||
// shape is the same either way.
|
||||
L.push(`plate 0 kind=poly slot_mouth=${r(take.slot.x)},${r(take.slot.y)} scale=1.00 rot=0 squash=1.00`);
|
||||
L.push('');
|
||||
for (const part of take.parts) {
|
||||
const bits = [`part ${part.name.padEnd(9)} kind=${part.kind} z=${part.z}`];
|
||||
if (part.color !== undefined) bits.push(`color=${part.color}`);
|
||||
if (part.kind === 'poly') bits.push('closed=1 fill=1');
|
||||
bits.push(`interp=${part.interp}`);
|
||||
if (part.parent) bits.push(`parent=${part.parent}`);
|
||||
L.push(bits.join(' '));
|
||||
}
|
||||
L.push('');
|
||||
for (const part of take.parts) {
|
||||
for (const k of part.keys) {
|
||||
if (k.hidden) { L.push(`key ${part.name.padEnd(9)} f=${k.f} hidden`); continue; }
|
||||
if (part.kind === 'plate') { L.push(`key ${part.name.padEnd(9)} f=${k.f} plate=${k.plate}`); continue; }
|
||||
const pts = k.pts.map((p) => `${r(p.x)},${r(p.y)}`).join(' ');
|
||||
// f is authoritative (what renders); src is the pre-snap extreme frame,
|
||||
// kept only as a tuning signal - a key dragged far means the minimum-hold
|
||||
// and the exposure grid are fighting.
|
||||
L.push(`key ${part.name.padEnd(9)} f=${k.f} src=${k.src} n=${k.pts.length} ${pts}`);
|
||||
}
|
||||
}
|
||||
L.push('');
|
||||
return L.join('\n');
|
||||
}
|
||||
20
selftest.html
Normal file
20
selftest.html
Normal file
|
|
@ -0,0 +1,20 @@
|
|||
<!doctype html>
|
||||
<html lang="en"><head><meta charset="utf-8"><title>selftest</title>
|
||||
<style>
|
||||
body{background:#0b0d13;color:#e8eaf0;font:13px/1.6 ui-monospace,Menlo,monospace;padding:20px}
|
||||
li{list-style:none} .p{color:#4ade80} .f{color:#f87171}
|
||||
b{font-weight:400;color:#8891a5} h1{font-size:14px}
|
||||
</style></head><body>
|
||||
<h1 id="head">running…</h1><ul id="out"></ul>
|
||||
<script type="module">
|
||||
import { run } from './js/selftest.js';
|
||||
let res;
|
||||
try { res = run(); }
|
||||
catch (e) { res = [{ name: 'harness threw: ' + e.message, pass: false, detail: String(e.stack).split('\n')[1] || '' }]; }
|
||||
const pass = res.filter(r => r.pass).length;
|
||||
const fail = res.length - pass;
|
||||
document.getElementById('head').textContent = `${fail === 0 ? 'PASS' : 'FAIL'} — ${pass}/${res.length} assertions`;
|
||||
document.title = `${fail === 0 ? 'PASS' : 'FAIL'} ${pass}/${res.length}`;
|
||||
document.getElementById('out').innerHTML = res.map(r =>
|
||||
`<li class="${r.pass ? 'p' : 'f'}">${r.pass ? '✓' : '✗'} ${r.name}${r.detail ? ` <b>${r.detail}</b>` : ''}</li>`).join('');
|
||||
</script></body></html>
|
||||
Loading…
Add table
Add a link
Reference in a new issue