roto: video -> take file builder with interactive tuning

Analysis half of the pipeline in docs/roto-puppet.md. Stabilises a face out
of a clip via a similarity fit on rigid landmarks, reduces the lip contour to
a fixed vertex budget, selects sparse keys on velocity minima, and previews
the result as flat indexed fills so timing can be judged without an Animator
Pro render.

- landmarks.js  ordered lip/oval rings; slot position is vertex identity
- mathutil.js   closed-form 2D similarity, Procrustes mean, transform smoothing
- pipeline.js   stabilise -> subsample -> key-select
- raster.js     indexed scanline fill, no antialiasing
- take.js       take-file writer
- selftest.js   29 assertions, incl. ring simplicity at every vertex budget

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Your Name 2026-09-24 14:38:07 -04:00
commit 7a6bdea157
13 changed files with 1277 additions and 0 deletions

3
.gitignore vendored Normal file
View file

@ -0,0 +1,3 @@
frames/
*.task
*.take

86
README.md Normal file
View file

@ -0,0 +1,86 @@
# roto
Video → **take file** builder for the Animator Pro rotoscope/puppet pipeline
described in `../docs/roto-puppet.md`.
This is the analysis and tuning half. It stabilises a face out of a clip, reduces
the lip contour to a handful of vertices, selects sparse keys on motion extremes,
and previews the result as flat indexed fills — so the *look and timing* can be
judged in seconds rather than through a minutes-long Animator Pro render. It
emits a `.take` file; nothing here touches Animator Pro.
All policy lives here. The take arrives at the renderer with its keys already
chosen.
## Run
```sh
python3 -m http.server 8777 # from this directory
# open http://127.0.0.1:8777
```
**Synthetic take** needs no video and exercises everything below detection.
For real footage:
```sh
./extract.sh /path/to/clip.mp4 24 # -> frames/0001.png …
```
then **Load frames**. MediaPipe's wasm is fetched from jsdelivr on first use;
`face_landmarker.task` is local.
Frames are pre-extracted rather than decoded in the page because browser video
seeking is approximate and `requestVideoFrameCallback` only delivers frames at
playback speed — neither gives a deterministic per-frame pass.
## Shooting for it
Near-frontal, good light, consistent scale, head reasonably still. Hold a neutral
closed mouth for a second at the top of the take: that frame is picked
automatically as the neutral and drives calibration and the placeholder plate.
Out-of-plane head rotation cannot be stabilised away by a 2D similarity
transform — the `residual` readout rises when it happens. The pipeline's answer
is hand-drawn head plates, which this tool does not yet do.
## Knobs
| Knob | What it does |
| --- | --- |
| vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. |
| min hold | Minimum frames between keys. |
| change gate | Mean vertex movement required before a new key is accepted. |
| vel smoothing | Window on the velocity signal used to find extremes. |
| anchor smoothing | Window on the four similarity parameters. Smooths the *transform*, never the contour. |
| exposure | Grid that key frames snap onto. |
| closed-mouth cut | Aperture below which the mouth interior is emitted as `hidden`. |
Keys go on **velocity minima**, not distance thresholds: a threshold fires at the
frame it was crossed — partway through a transition — so poses land mushy and
late. The timeline shows the velocity curve, candidate minima, accepted keys
(`f`) and their pre-snap extremes (`src`); a large `f`/`src` gap means min-hold
and exposure are fighting.
## Tests
```sh
chromium --headless --virtual-time-budget=8000 --dump-dom \
http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+'
```
Or open `selftest.html`. 29 assertions over the stages below detection.
The ring-simplicity check is the load-bearing one. Because `hold` parts *cut*
between poses instead of interpolating, a ring whose vertex order is wrong
self-intersects and renders as blocks meeting at corners — and it is invisible at
odd `verts/2` and obvious at even, so it needs an assertion rather than an
eyeball.
## Not done yet
Eyes and irises; hand-drawn head plates and per-plate mouth slots; real
performer→character calibration (currently identity, fitting the face oval to the
canvas); the override layer; anything on the Animator Pro side. The placeholder
plate is a frozen face-oval polygon — it exists so the mouth has a face to read
against, not to look good.

17
extract.sh Executable file
View file

@ -0,0 +1,17 @@
#!/usr/bin/env bash
# Extract a clip to a PNG sequence for the take builder.
#
# Frames are pre-extracted rather than decoded in the page on purpose: browser
# video seeking by currentTime is approximate and requestVideoFrameCallback only
# delivers frames at playback speed, so neither gives a deterministic per-frame
# pass. A PNG sequence is exact, instantly seekable, and reproducible.
set -euo pipefail
src="${1:?usage: ./extract.sh CLIP [FPS] [OUTDIR]}"
fps="${2:-24}"
out="${3:-frames}"
rm -rf "$out"
mkdir -p "$out"
ffmpeg -hide_banner -loglevel warning -i "$src" -vf "fps=$fps" "$out/%04d.png"
echo "$(ls -1 "$out" | wc -l) frames at ${fps}fps -> $out/"

123
index.html Normal file
View file

@ -0,0 +1,123 @@
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>roto — take builder</title>
<style>
:root {
--bg: #0b0d13; --panel: #12151f; --line: #232836; --fg: #e8eaf0;
--dim: #8891a5; --accent: #fbbf24; --ok: #4ade80; --err: #f87171;
}
* { box-sizing: border-box; }
body { margin: 0; background: var(--bg); color: var(--fg);
font: 13px/1.5 ui-monospace, SFMono-Regular, Menlo, monospace; }
header { padding: 14px 18px; border-bottom: 1px solid var(--line);
display: flex; gap: 14px; align-items: baseline; flex-wrap: wrap; }
h1 { font-size: 14px; margin: 0; letter-spacing: .04em; }
h1 span { color: var(--dim); font-weight: 400; }
main { padding: 18px; display: flex; flex-direction: column; gap: 18px; }
.row { display: flex; gap: 18px; flex-wrap: wrap; }
.panel { background: var(--panel); border: 1px solid var(--line); border-radius: 4px; padding: 12px; }
.panel h2 { font-size: 11px; text-transform: uppercase; letter-spacing: .08em;
color: var(--dim); margin: 0 0 8px; font-weight: 500; }
canvas { display: block; image-rendering: pixelated; max-width: 100%; border-radius: 2px; }
#cv-timeline { width: 100%; height: 90px; }
button, input[type=text] {
font: inherit; background: #1c2130; color: var(--fg);
border: 1px solid var(--line); border-radius: 3px; padding: 5px 11px; cursor: pointer; }
button:hover { border-color: var(--accent); }
input[type=text] { cursor: text; }
label.ctl { display: grid; grid-template-columns: 116px 1fr 44px;
gap: 10px; align-items: center; margin-bottom: 6px; }
label.ctl span:first-child { color: var(--dim); }
label.ctl output { text-align: right; color: var(--accent); }
input[type=range] { width: 100%; accent-color: var(--accent); }
#status { color: var(--dim); }
#status.ok { color: var(--ok); } #status.err { color: var(--err); } #status.warn { color: var(--accent); }
#readout { color: var(--dim); font-size: 12px; }
#sheet { display: flex; flex-wrap: wrap; gap: 8px; }
#sheet .cell { display: flex; flex-direction: column; gap: 3px; cursor: pointer; }
#sheet .cell span { color: var(--dim); font-size: 11px; }
#sheet .cell:hover span { color: var(--accent); }
#palette { display: flex; gap: 12px; flex-wrap: wrap; }
.sw { display: flex; align-items: center; gap: 5px; color: var(--dim); font-size: 11px; }
.sw input { width: 26px; height: 20px; padding: 0; border: 1px solid var(--line); background: none; }
.legend { color: var(--dim); font-size: 11px; margin-top: 6px; }
.legend b { font-weight: 400; }
.k { color: var(--accent); } .s { color: var(--err); } .c { color: #5b6478; }
</style>
</head>
<body>
<header>
<h1>roto <span>— video → take builder</span></h1>
<button id="btn-synth">Synthetic take</button>
<input type="text" id="framedir" value="frames" size="8" title="frame directory">
<button id="btn-frames">Load frames</button>
<button id="btn-play">Play</button>
<input type="text" id="takename" value="line_01" size="10" title="take name">
<button id="btn-export">Export .take</button>
<span id="status"></span>
</header>
<main>
<div class="row">
<div class="panel">
<h2>source + landmarks</h2>
<canvas id="cv-source"></canvas>
<div class="legend">outer lip <b style="color:#4ade80">—</b> · inner lip <b style="color:#f87171">—</b></div>
</div>
<div class="panel">
<h2>stabilised (head-local)</h2>
<canvas id="cv-stab"></canvas>
<div class="legend">live contour · <b class="k">—</b> active key pose · should sit still except the mouth</div>
</div>
<div class="panel">
<h2>flat render — 320×200 indexed</h2>
<canvas id="cv-render"></canvas>
<div class="legend" id="framelabel"></div>
</div>
</div>
<div class="row">
<div class="panel" style="flex:1 1 380px">
<h2>knobs</h2>
<label class="ctl"><span>vertices</span><input type="range" id="verts" min="4" max="16" step="2" value="8"><output id="vertsv"></output></label>
<label class="ctl"><span>min hold</span><input type="range" id="minHold" min="1" max="8" value="2"><output id="minHoldv"></output></label>
<label class="ctl"><span>change gate</span><input type="range" id="distThresh" min="0" max="40" value="6"><output id="distThreshv"></output></label>
<label class="ctl"><span>vel smoothing</span><input type="range" id="velSmooth" min="1" max="11" step="2" value="3"><output id="velSmoothv"></output></label>
<label class="ctl"><span>anchor smoothing</span><input type="range" id="smoothWin" min="1" max="21" step="2" value="5"><output id="smoothWinv"></output></label>
<label class="ctl"><span>exposure</span><input type="range" id="exposure" min="1" max="4" value="2"><output id="exposurev"></output></label>
<label class="ctl"><span>closed-mouth cut</span><input type="range" id="apertureThresh" min="0" max="400" value="120"><output id="apertureThreshv"></output></label>
<div class="legend" id="readout"></div>
</div>
<div class="panel" style="flex:1 1 300px">
<h2>palette</h2>
<div id="palette"></div>
<div class="legend" style="margin-top:12px">
Flat indexed fills, no antialiasing — the rasteriser writes palette
indices, the way <code>csd_render_poly</code> does.
</div>
</div>
</div>
<div class="panel">
<h2>timeline</h2>
<canvas id="cv-timeline"></canvas>
<input type="range" id="scrub" min="0" max="0" value="0" style="width:100%;margin-top:8px">
<div class="legend">
velocity curve · <b class="c">|</b> candidate minima ·
<b class="k">|</b> key as rendered (<code>f</code>) ·
<b class="s">|</b> true extreme (<code>src</code>) — a large gap means min-hold and exposure are fighting
</div>
</div>
<div class="panel">
<h2>selected poses — the vocabulary the take actually contains</h2>
<div id="sheet"></div>
</div>
</main>
<script type="module" src="./js/app.js"></script>
</body>
</html>

403
js/app.js Normal file
View file

@ -0,0 +1,403 @@
import { FaceLandmarker, FilesetResolver } from 'https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@1.0.1/vision_bundle.mjs';
import { LIPS_OUTER, LIPS_INNER, FACE_OVAL } from './landmarks.js';
import { stabilize, toRasterRing, selectKeys, activeKey } from './pipeline.js';
import { IndexedRaster } from './raster.js';
import { writeTake } from './take.js';
import { synthDense } from './synth.js';
const RW = 320, RH = 200, ZOOM = 2;
const PALETTE = [
{ name: 'bg', hex: '#12141c' },
{ name: 'skin_base', hex: '#b07a5a' },
{ name: 'skin_dark', hex: '#7a4f3a' },
{ name: 'mouth_dark', hex: '#24161a' },
{ name: 'skin_lite', hex: '#d9a884' },
];
const IDX = { bg: 0, base: 1, dark: 2, mouth: 3, lite: 4 };
const state = {
dense: null, images: [], stab: null, xform: null,
plate: null, neutral: 0, keysOuter: null, keysInner: null,
frame: 0, playing: false, source: 'none',
};
const el = (id) => document.getElementById(id);
const opts = () => ({
verts: +el('verts').value,
minHold: +el('minHold').value,
distThresh: +el('distThresh').value / 10,
velSmooth: +el('velSmooth').value,
smoothWin: +el('smoothWin').value,
exposure: +el('exposure').value,
apertureThresh: +el('apertureThresh').value / 1000,
});
function status(msg, kind = '') {
const s = el('status');
s.textContent = msg;
s.className = kind;
}
/* ---------- frame loading ---------- */
function loadImage(src) {
return new Promise((res) => {
const im = new Image();
im.onload = () => res(im);
im.onerror = () => res(null);
im.src = src;
});
}
async function loadFrameSequence() {
const dir = el('framedir').value.replace(/\/$/, '');
const imgs = [];
for (let i = 1; i <= 900; i++) {
const name = `${dir}/${String(i).padStart(4, '0')}.png`;
const im = await loadImage(name);
if (!im) break;
imgs.push(im);
if (i % 10 === 0) status(`loading frames… ${i}`);
}
return imgs;
}
/* ---------- detection ---------- */
let landmarker = null;
async function initLandmarker() {
if (landmarker) return landmarker;
status('loading MediaPipe wasm…');
const fileset = await FilesetResolver.forVisionTasks(
'https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@1.0.1/wasm');
landmarker = await FaceLandmarker.createFromOptions(fileset, {
baseOptions: { modelAssetPath: './face_landmarker.task', delegate: 'GPU' },
runningMode: 'IMAGE',
numFaces: 1,
});
return landmarker;
}
async function detectAll(images) {
const lm = await initLandmarker();
const cv = document.createElement('canvas');
const dense = [];
const missing = [];
for (let i = 0; i < images.length; i++) {
const im = images[i];
cv.width = im.naturalWidth; cv.height = im.naturalHeight;
cv.getContext('2d').drawImage(im, 0, 0);
const out = lm.detect(cv);
if (out.faceLandmarks && out.faceLandmarks.length) {
dense.push(out.faceLandmarks[0]);
} else {
// Hold the previous frame rather than dropping it, so frame indices stay
// aligned with the source sequence. A gap is reported, not hidden.
missing.push(i);
dense.push(dense.length ? dense[dense.length - 1] : null);
}
if (i % 5 === 0) status(`detecting… ${i + 1}/${images.length}`);
}
if (dense[0] === null) throw new Error('no face found in the first frame');
return { dense, missing };
}
/* ---------- calibration ---------- */
// v1 calibration: fit the reference face oval's bounding box to a target box on
// the character raster. Identity retargeting - it makes any clip frame sensibly,
// but a real project replaces this with a measured performer->character map.
function makeXform(stab) {
const oval = stab.oval[state.neutral];
let x0 = Infinity, y0 = Infinity, x1 = -Infinity, y1 = -Infinity;
for (const p of oval) {
x0 = Math.min(x0, p.x); y0 = Math.min(y0, p.y);
x1 = Math.max(x1, p.x); y1 = Math.max(y1, p.y);
}
const targetH = RH * 0.80;
const s = targetH / (y1 - y0);
const cx = (x0 + x1) / 2, cy = (y0 + y1) / 2;
return (p) => ({ x: (p.x - cx) * s + RW / 2, y: (p.y - cy) * s + RH / 2 });
}
/* ---------- build ---------- */
function rebuild() {
if (!state.dense) return;
const o = opts();
state.stab = stabilize(state.dense, o.smoothWin);
// Neutral frame = most closed mouth in the first quarter of the shot, which is
// where the performer is asked to hold a neutral closed mouth.
const ap = state.stab.aperture;
const head = Math.max(1, Math.floor(ap.length / 4));
let best = 0;
for (let i = 0; i < head; i++) if (ap[i] < ap[best]) best = i;
state.neutral = best;
state.xform = makeXform(state.stab);
state.plate = state.stab.oval[state.neutral].map(state.xform);
const outer = state.stab.outer.map((r) => toRasterRing(r, LIPS_OUTER, o.verts, state.xform));
const inner = state.stab.inner.map((r) => toRasterRing(r, LIPS_INNER, o.verts, state.xform));
state.shapesOuter = outer;
state.shapesInner = inner;
state.keysOuter = selectKeys(outer, o);
// The interior is a child of the lip silhouette: it keys on exactly its
// parent's frames, never independently, or it swims.
state.keysInner = {
...state.keysOuter,
keys: state.keysOuter.keys.map((k) => ({ ...k })),
};
const apMax = Math.max(...ap);
state.hidden = state.keysOuter.keys.map((k) => ap[k.src] / apMax < o.apertureThresh);
el('readout').textContent =
`${state.dense.length} frames · ${state.keysOuter.candidates.length} candidates → ` +
`${state.keysOuter.keys.length} keys · ${o.verts} verts · neutral f${state.neutral} · ` +
`residual ${(state.stab.residual.reduce((a, b) => a + b, 0) / state.dense.length).toFixed(4)}`;
drawTimeline();
drawContactSheet();
draw();
}
/* ---------- drawing ---------- */
function renderFrame(f) {
const r = new IndexedRaster(RW, RH);
r.clear(IDX.bg);
if (state.plate) r.fillPoly(state.plate, IDX.base);
const ko = activeKey(state.keysOuter.keys, f);
const ki = activeKey(state.keysInner.keys, f);
if (ko) r.fillPoly(state.shapesOuter[ko.src], IDX.dark);
if (ki) {
const slot = state.keysOuter.keys.indexOf(ki);
if (!state.hidden[slot]) r.fillPoly(state.shapesInner[ki.src], IDX.mouth);
}
return r;
}
function blit(canvas, raster, zoom) {
canvas.width = RW * zoom; canvas.height = RH * zoom;
canvas.getContext('2d').putImageData(raster.toImageData(PALETTE.map((p) => p.hex), zoom), 0, 0);
}
function draw() {
const f = state.frame;
el('framelabel').textContent =
`f ${String(f).padStart(3)} / ${state.dense.length - 1} ` +
`key ${activeKey(state.keysOuter.keys, f).f}`;
// pane 1: source with the raw contour overlaid
const c1 = el('cv-source'), g1 = c1.getContext('2d');
c1.width = RW * ZOOM; c1.height = RH * ZOOM;
g1.fillStyle = '#000'; g1.fillRect(0, 0, c1.width, c1.height);
const im = state.images[f];
if (im) {
const s = Math.min(c1.width / im.naturalWidth, c1.height / im.naturalHeight);
const w = im.naturalWidth * s, h = im.naturalHeight * s;
g1.drawImage(im, (c1.width - w) / 2, (c1.height - h) / 2, w, h);
g1.save();
g1.translate((c1.width - w) / 2, (c1.height - h) / 2);
strokeRing(g1, LIPS_OUTER.map((i) => state.dense[f][i]), w, h, '#4ade80');
strokeRing(g1, LIPS_INNER.map((i) => state.dense[f][i]), w, h, '#f87171');
g1.restore();
} else {
g1.fillStyle = '#555'; g1.font = '13px system-ui';
g1.fillText('synthetic — no source frames', 14, 24);
strokeRing(g1, LIPS_OUTER.map((i) => state.dense[f][i]), c1.width, c1.height, '#4ade80');
strokeRing(g1, LIPS_INNER.map((i) => state.dense[f][i]), c1.width, c1.height, '#f87171');
}
// pane 2: stabilised contour against a fixed reference cross.
// If stabilisation works, this contour stays put except for mouth motion.
const c2 = el('cv-stab'), g2 = c2.getContext('2d');
c2.width = RW * ZOOM; c2.height = RH * ZOOM;
g2.fillStyle = '#0d0f16'; g2.fillRect(0, 0, c2.width, c2.height);
g2.strokeStyle = '#2a2f3e'; g2.lineWidth = 1;
g2.beginPath();
g2.moveTo(c2.width / 2, 0); g2.lineTo(c2.width / 2, c2.height);
g2.moveTo(0, c2.height / 2); g2.lineTo(c2.width, c2.height / 2);
g2.stroke();
const pl = (pts) => pts.map((p) => { const q = state.xform(p); return { x: q.x * ZOOM, y: q.y * ZOOM }; });
strokePts(g2, pl(state.stab.oval[f]), '#3b4a63');
strokePts(g2, pl(state.stab.outer[f]), '#4ade80');
strokePts(g2, pl(state.stab.inner[f]), '#f87171');
const ko = activeKey(state.keysOuter.keys, f);
strokePts(g2, state.shapesOuter[ko.src].map((p) => ({ x: p.x * ZOOM, y: p.y * ZOOM })), '#fbbf24', 2);
// pane 3: the flat indexed render
blit(el('cv-render'), renderFrame(f), ZOOM);
}
function strokeRing(g, pts, w, h, color) {
strokePts(g, pts.map((p) => ({ x: p.x * w, y: p.y * h })), color);
}
function strokePts(g, pts, color, lw = 1) {
g.strokeStyle = color; g.lineWidth = lw;
g.beginPath();
pts.forEach((p, i) => (i ? g.lineTo(p.x, p.y) : g.moveTo(p.x, p.y)));
g.closePath(); g.stroke();
}
function drawTimeline() {
const c = el('cv-timeline'), g = c.getContext('2d');
const N = state.dense.length;
const W = Math.max(640, N * 8), H = 90;
c.width = W; c.height = H;
const px = W / N;
g.fillStyle = '#0d0f16'; g.fillRect(0, 0, W, H);
const vel = state.keysOuter.velocity;
const vmax = Math.max(...vel) || 1;
g.strokeStyle = '#3b4a63'; g.lineWidth = 1; g.beginPath();
vel.forEach((v, i) => {
const x = i * px + px / 2, y = H - 18 - (v / vmax) * (H - 34);
i ? g.lineTo(x, y) : g.moveTo(x, y);
});
g.stroke();
g.fillStyle = '#5b6478';
for (const t of state.keysOuter.candidates) g.fillRect(t * px + px / 2 - 0.5, H - 18, 1, 6);
for (const k of state.keysOuter.keys) {
g.fillStyle = '#fbbf24';
g.fillRect(k.f * px + px / 2 - 1, 6, 2, H - 24); // f: what renders
if (k.src !== k.f) {
g.fillStyle = '#f87171';
g.fillRect(k.src * px + px / 2 - 0.5, 6, 1, 10); // src: the true extreme
}
}
g.fillStyle = '#e8eaf0';
g.fillRect(state.frame * px + px / 2 - 0.5, 0, 1, H);
}
function drawContactSheet() {
const host = el('sheet');
host.innerHTML = '';
state.keysOuter.keys.forEach((k, i) => {
const cell = document.createElement('div');
cell.className = 'cell';
const cv = document.createElement('canvas');
blit(cv, renderFrame(k.f), 1);
cv.style.width = '104px';
const cap = document.createElement('span');
cap.textContent = `f${k.f}${k.src !== k.f ? ` ←${k.src}` : ''}${state.hidden[i] ? ' ·closed' : ''}`;
cell.append(cv, cap);
cell.onclick = () => { state.frame = k.f; draw(); drawTimeline(); };
host.append(cell);
});
}
/* ---------- export ---------- */
function exportTake() {
const keys = state.keysOuter.keys;
const take = {
name: el('takename').value || 'line_01',
frames: state.dense.length, width: RW, height: RH,
exposure: opts().exposure,
palette: PALETTE,
slot: { x: RW / 2, y: RH / 2 },
parts: [
{ name: 'head', kind: 'plate', z: 0, interp: 'hold', keys: [{ f: 0, plate: 0 }] },
{ name: 'mouth', kind: 'poly', z: 30, color: 'skin_dark', interp: 'hold',
keys: keys.map((k) => ({ f: k.f, src: k.src, pts: state.shapesOuter[k.src] })) },
{ name: 'mouth_in', kind: 'poly', z: 31, color: 'mouth_dark', interp: 'hold', parent: 'mouth',
keys: keys.map((k, i) => state.hidden[i]
? { f: k.f, hidden: true }
: { f: k.f, src: k.src, pts: state.shapesInner[k.src] }) },
],
};
const text = writeTake(take);
const a = document.createElement('a');
a.href = URL.createObjectURL(new Blob([text], { type: 'text/plain' }));
a.download = `${take.name}.take`;
a.click();
status(`exported ${take.name}.take — ${keys.length} keys`, 'ok');
}
/* ---------- wiring ---------- */
async function runFrames() {
try {
status('loading frames…');
const images = await loadFrameSequence();
if (!images.length) {
status(`no frames found in ${el('framedir').value}/ — run extract.sh first`, 'err');
return;
}
const { dense, missing } = await detectAll(images);
state.images = images; state.dense = dense; state.source = 'video';
el('scrub').max = dense.length - 1;
state.frame = 0;
rebuild();
status(missing.length
? `${images.length} frames · no face on ${missing.length} (held previous)`
: `${images.length} frames detected`, missing.length ? 'warn' : 'ok');
} catch (e) {
status(`${e.message}`, 'err');
console.error(e);
}
}
function runSynthetic() {
state.images = [];
state.dense = synthDense(72);
state.source = 'synthetic';
el('scrub').max = 71;
state.frame = 0;
rebuild();
status('synthetic take — no video needed; exercises the whole chain below detection', 'ok');
}
for (const id of ['verts','minHold','distThresh','velSmooth','smoothWin','exposure','apertureThresh']) {
el(id).addEventListener('input', () => {
el(id + 'v').textContent = id === 'distThresh' ? (+el(id).value / 10).toFixed(1)
: id === 'apertureThresh' ? (+el(id).value / 1000).toFixed(3)
: el(id).value;
rebuild();
});
el(id + 'v').textContent = el(id).value;
}
el('scrub').addEventListener('input', (e) => { state.frame = +e.target.value; draw(); drawTimeline(); });
el('btn-frames').onclick = runFrames;
el('btn-synth').onclick = runSynthetic;
el('btn-export').onclick = exportTake;
el('btn-play').onclick = () => {
state.playing = !state.playing;
el('btn-play').textContent = state.playing ? 'Stop' : 'Play';
if (state.playing) tick();
};
let last = 0;
function tick(ts = 0) {
if (!state.playing) return;
if (ts - last > 1000 / 24) {
last = ts;
state.frame = (state.frame + 1) % state.dense.length;
el('scrub').value = state.frame;
draw(); drawTimeline();
}
requestAnimationFrame(tick);
}
PALETTE.forEach((p) => {
const sw = document.createElement('label');
sw.className = 'sw';
const inp = document.createElement('input');
inp.type = 'color'; inp.value = p.hex;
inp.oninput = () => { p.hex = inp.value; if (state.dense) { draw(); drawContactSheet(); } };
sw.append(inp, document.createTextNode(p.name));
el('palette').append(sw);
});
status('ready — "Synthetic take" works with no video; "Load frames" reads ./frames/');

54
js/landmarks.js Normal file
View file

@ -0,0 +1,54 @@
// MediaPipe FaceLandmarker index tables.
// Ring arrays are ORDERED traversals, not raw connection sets: vertex position
// within a ring is the vertex's identity, and every downstream stage depends on
// that ordering being stable. See docs/roto-puppet.md, "Fixed topology".
// Rigid landmarks for the similarity fit. Eye corners, nose bridge, nose tip.
// Nothing here may be a feature that moves under performance: including the
// mouth or brows bleeds performance into the stabilization.
export const RIGID = [33, 133, 362, 263, 168, 6, 1];
// Outer lip ring, clockwise from the right corner over the top.
// index 0 = right corner, 5 = top centre, 10 = left corner, 15 = bottom centre.
export const LIPS_OUTER = [
61, 185, 40, 39, 37, 0, 267, 269, 270, 409,
291, 375, 321, 405, 314, 17, 84, 181, 91, 146,
];
// Inner lip ring, same orientation and the same four cardinal positions.
export const LIPS_INNER = [
78, 191, 80, 81, 82, 13, 312, 311, 310, 415,
308, 324, 318, 402, 317, 14, 87, 178, 88, 95,
];
// Inner upper / lower lip centres. Their separation is the aperture signal that
// decides whether the mouth interior is present at all.
export const APERTURE = [13, 14];
// Face oval, used only to derive the placeholder plate in v1.
export const FACE_OVAL = [
10, 338, 297, 332, 284, 251, 389, 356, 454, 323, 361, 288,
397, 365, 379, 378, 400, 377, 152, 148, 176, 149, 150, 136,
172, 58, 132, 93, 234, 127, 162, 21, 54, 103, 67, 109,
];
// Eye corners, for the calibration box and for reporting fit residual.
export const EYE_INNER = [133, 362];
// Pick `n` slots from a ring of `len` by even spacing. Returns RING POSITIONS,
// not landmark ids: positions are the vertex identity downstream, and mapping ids
// back to positions with indexOf would silently pick the wrong slot if a table
// ever repeated an id.
//
// For even n this naturally lands on the cardinal positions (corners and lip
// centres) of a 20-point ring. Fixed indices, never adaptive decimation: the
// vertex at slot k means the same thing on every frame of the shot.
export function subsampleSlots(len, n) {
const out = [];
for (let k = 0; k < n; k++) out.push(Math.round((k * len) / n) % len);
return out;
}
export function subsampleRing(ring, n) {
return subsampleSlots(ring.length, n).map((s) => ring[s]);
}

107
js/mathutil.js Normal file
View file

@ -0,0 +1,107 @@
// 2D similarity transforms and temporal smoothing.
// Least-squares similarity (translation + rotation + uniform scale, 4 DOF)
// mapping P onto Q. Closed form; no iteration.
//
// Deliberately NOT affine or homography: the extra degrees of freedom absorb
// out-of-plane head rotation as shear/perspective and smear it into the mouth.
// Four DOF removes exactly translation, roll and depth-scale, and leaves yaw and
// pitch as a measurable residual.
export function fitSimilarity(P, Q) {
const n = P.length;
let pcx = 0, pcy = 0, qcx = 0, qcy = 0;
for (let i = 0; i < n; i++) {
pcx += P[i].x; pcy += P[i].y;
qcx += Q[i].x; qcy += Q[i].y;
}
pcx /= n; pcy /= n; qcx /= n; qcy /= n;
let a = 0, b = 0, norm = 0;
for (let i = 0; i < n; i++) {
const px = P[i].x - pcx, py = P[i].y - pcy;
const qx = Q[i].x - qcx, qy = Q[i].y - qcy;
a += px * qx + py * qy; // dot
b += px * qy - py * qx; // cross
norm += px * px + py * py;
}
const theta = Math.atan2(b, a);
const s = norm > 1e-12 ? Math.hypot(a, b) / norm : 1;
const c = Math.cos(theta), sn = Math.sin(theta);
return {
s, theta,
tx: qcx - s * (c * pcx - sn * pcy),
ty: qcy - s * (sn * pcx + c * pcy),
};
}
export function applySim(tf, p) {
const c = Math.cos(tf.theta), sn = Math.sin(tf.theta);
return {
x: tf.s * (c * p.x - sn * p.y) + tf.tx,
y: tf.s * (sn * p.x + c * p.y) + tf.ty,
};
}
export function applySimAll(tf, pts) {
return pts.map((p) => applySim(tf, p));
}
// Residual RMS after the fit, in the units of Q. Rises with out-of-plane
// rotation, so it is the signal for "this section is not stabilisable".
export function fitResidual(tf, P, Q) {
let acc = 0;
for (let i = 0; i < P.length; i++) {
const m = applySim(tf, P[i]);
acc += (m.x - Q[i].x) ** 2 + (m.y - Q[i].y) ** 2;
}
return Math.sqrt(acc / P.length);
}
// Generalised Procrustes: the reference is the MEAN rigid configuration over the
// shot, not frame zero, so no single frame's idiosyncrasies get baked into every
// other frame. Three passes is plenty.
export function procrustesMean(framesRigid, iters = 3) {
let ref = framesRigid[0].map((p) => ({ x: p.x, y: p.y }));
for (let it = 0; it < iters; it++) {
const acc = ref.map(() => ({ x: 0, y: 0 }));
for (const rig of framesRigid) {
const tf = fitSimilarity(rig, ref);
const moved = applySimAll(tf, rig);
for (let i = 0; i < acc.length; i++) { acc[i].x += moved[i].x; acc[i].y += moved[i].y; }
}
ref = acc.map((p) => ({ x: p.x / framesRigid.length, y: p.y / framesRigid.length }));
}
return ref;
}
function movingAverage(vals, win) {
if (win <= 1) return vals.slice();
const half = Math.floor(win / 2), out = new Array(vals.length);
for (let i = 0; i < vals.length; i++) {
let acc = 0, cnt = 0;
for (let j = i - half; j <= i + half; j++) {
const k = Math.min(vals.length - 1, Math.max(0, j));
acc += vals[k]; cnt++;
}
out[i] = acc / cnt;
}
return out;
}
export { movingAverage };
// Smooth the four transform parameters, NEVER the contour. Landmark jitter of a
// pixel is smeared into the mouth by the inverse transform, so the transform is
// where the low-pass belongs; smoothing the contour would destroy the
// performance, which is the entire asset.
// Angles are smoothed as (cos, sin) so wrapping cannot produce a spike.
export function smoothTransforms(tfs, win) {
const c = movingAverage(tfs.map((t) => Math.cos(t.theta)), win);
const sn = movingAverage(tfs.map((t) => Math.sin(t.theta)), win);
const s = movingAverage(tfs.map((t) => t.s), win);
const tx = movingAverage(tfs.map((t) => t.tx), win);
const ty = movingAverage(tfs.map((t) => t.ty), win);
return tfs.map((_, i) => ({
theta: Math.atan2(sn[i], c[i]), s: s[i], tx: tx[i], ty: ty[i],
}));
}

108
js/pipeline.js Normal file
View file

@ -0,0 +1,108 @@
// Analysis: dense track -> stabilised head-local contours -> selected keys.
// All policy lives here, never in the renderer. See docs/roto-puppet.md,
// "The take is the contract".
import { RIGID, LIPS_OUTER, LIPS_INNER, APERTURE, FACE_OVAL, EYE_INNER, subsampleSlots } from './landmarks.js';
import { fitSimilarity, applySimAll, applySim, fitResidual, procrustesMean, smoothTransforms, movingAverage } from './mathutil.js';
const pick = (lm, idx) => idx.map((i) => ({ x: lm[i].x, y: lm[i].y }));
// Stage 1-3: fit the rigid transform per frame, smooth its parameters, then map
// every contour through it into the reference frame. The result is head-local:
// translation, roll and depth-scale of the head are gone.
export function stabilize(dense, smoothWin) {
const rigid = dense.map((f) => pick(f, RIGID));
const ref = procrustesMean(rigid);
const raw = rigid.map((r) => fitSimilarity(r, ref));
const tfs = smoothTransforms(raw, smoothWin);
return {
ref,
transforms: tfs,
// Residual rises with out-of-plane rotation, which no 2D similarity can
// remove. High values mean this section wants a different head plate.
residual: tfs.map((tf, i) => fitResidual(tf, rigid[i], ref)),
outer: dense.map((f, i) => applySimAll(tfs[i], pick(f, LIPS_OUTER))),
inner: dense.map((f, i) => applySimAll(tfs[i], pick(f, LIPS_INNER))),
oval: dense.map((f, i) => applySimAll(tfs[i], pick(f, FACE_OVAL))),
eyes: dense.map((f, i) => applySimAll(tfs[i], pick(f, EYE_INNER))),
aperture: dense.map((f, i) => {
const a = applySimAll(tfs[i], pick(f, APERTURE));
return Math.hypot(a[0].x - a[1].x, a[0].y - a[1].y);
}),
};
}
// Stage 4: fixed-index subsample of a stabilised ring, then map from normalised
// face space into character raster space.
export function toRasterRing(stabRing, ringTable, n, xform) {
return subsampleSlots(ringTable.length, n).map((s) => xform(stabRing[s]));
}
// Stage 6: key selection.
//
// Keys go on velocity MINIMA, not on distance thresholds. A threshold fires at
// the frame it was crossed - partway through a transition - so every pose lands
// mushy and late. A minimum is where the shape is momentarily parked, which is
// the pose a viewer actually reads.
//
// Minima alone are not enough: during a long hold the velocity wobbles near zero
// and produces a key per wobble. So a candidate minimum is only accepted if the
// shape has actually moved since the last accepted key (distThresh) and the
// minimum hold has elapsed (minHold).
export function selectKeys(shapes, opts) {
const { minHold, distThresh, velSmooth, exposure } = opts;
const N = shapes.length;
if (N === 0) return { keys: [], velocity: [], candidates: [] };
const vel = new Array(N).fill(0);
for (let t = 1; t < N; t++) {
let acc = 0;
for (let i = 0; i < shapes[t].length; i++) {
acc += Math.hypot(shapes[t][i].x - shapes[t - 1][i].x, shapes[t][i].y - shapes[t - 1][i].y);
}
vel[t] = acc / shapes[t].length;
}
const sv = movingAverage(vel, velSmooth);
const candidates = [];
for (let t = 1; t < N - 1; t++) {
if (sv[t] <= sv[t - 1] && sv[t] <= sv[t + 1]) candidates.push(t);
}
const shapeDist = (a, b) => {
let acc = 0;
for (let i = 0; i < a.length; i++) acc += Math.hypot(a[i].x - b[i].x, a[i].y - b[i].y);
return acc / a.length;
};
const accepted = [0];
for (const t of candidates) {
const last = accepted[accepted.length - 1];
if (t - last < minHold) continue;
if (shapeDist(shapes[t], shapes[last]) < distThresh) continue;
accepted.push(t);
}
// Snap onto the exposure grid. f is what renders; src is provenance.
const keys = [];
for (const src of accepted) {
const f = Math.round(src / exposure) * exposure;
const prev = keys[keys.length - 1];
if (prev && prev.f === f) {
// Two extremes collapsed onto one grid slot: keep the stronger one.
if (sv[src] < sv[prev.src]) { prev.src = src; prev.frame = src; }
continue;
}
keys.push({ f, src, frame: src });
}
return { keys, velocity: sv, candidates };
}
// Resolve which key is live on a given output frame under interp=hold.
// "Most recent key at or before f" - lookup, not policy.
export function activeKey(keys, f) {
let hit = keys[0];
for (const k of keys) { if (k.f <= f) hit = k; else break; }
return hit;
}

83
js/raster.js Normal file
View file

@ -0,0 +1,83 @@
// Indexed flat-fill rasteriser.
//
// Canvas2D antialiases path fills, and antialiasing is exactly what the target
// idiom does not have: Animator Pro fills polygons into a 256-colour indexed
// raster with hard edges (csd_render_poly). A preview that antialiases would
// misrepresent the look it exists to judge, so this writes palette indices into
// a byte buffer with an even-odd scanline fill and expands to RGBA only at the
// very end.
export class IndexedRaster {
constructor(w, h) {
this.w = w; this.h = h;
this.buf = new Uint8Array(w * h);
}
clear(index) { this.buf.fill(index); }
// Even-odd scanline fill. Samples at pixel centres (y + 0.5), so a polygon
// edge landing exactly on a pixel boundary resolves consistently.
fillPoly(pts, index) {
const n = pts.length;
if (n < 3) return;
let minY = Infinity, maxY = -Infinity;
for (const p of pts) { if (p.y < minY) minY = p.y; if (p.y > maxY) maxY = p.y; }
const y0 = Math.max(0, Math.ceil(minY - 0.5));
const y1 = Math.min(this.h - 1, Math.floor(maxY - 0.5) + 1);
const xs = [];
for (let y = y0; y <= y1; y++) {
const sy = y + 0.5;
xs.length = 0;
for (let i = 0; i < n; i++) {
const a = pts[i], b = pts[(i + 1) % n];
if (a.y === b.y) continue;
const lo = Math.min(a.y, b.y), hi = Math.max(a.y, b.y);
if (sy < lo || sy >= hi) continue;
xs.push(a.x + ((sy - a.y) / (b.y - a.y)) * (b.x - a.x));
}
if (xs.length < 2) continue;
xs.sort((p, q) => p - q);
for (let k = 0; k + 1 < xs.length; k += 2) {
const xa = Math.max(0, Math.ceil(xs[k] - 0.5));
const xb = Math.min(this.w - 1, Math.floor(xs[k + 1] - 0.5));
const row = y * this.w;
for (let x = xa; x <= xb; x++) this.buf[row + x] = index;
}
}
}
fillDisc(cx, cy, r, index) {
const rr = r * r;
const y0 = Math.max(0, Math.floor(cy - r)), y1 = Math.min(this.h - 1, Math.ceil(cy + r));
const x0 = Math.max(0, Math.floor(cx - r)), x1 = Math.min(this.w - 1, Math.ceil(cx + r));
for (let y = y0; y <= y1; y++) {
for (let x = x0; x <= x1; x++) {
const dx = x + 0.5 - cx, dy = y + 0.5 - cy;
if (dx * dx + dy * dy <= rr) this.buf[y * this.w + x] = index;
}
}
}
// Expand indices through the palette into an ImageData at integer zoom.
// Nearest-neighbour by construction, so no filtering softens the result.
toImageData(palette, zoom = 1) {
const W = this.w * zoom, H = this.h * zoom;
const img = new ImageData(W, H);
const d = img.data;
const rgb = palette.map(hexToRgb);
for (let y = 0; y < H; y++) {
const srow = Math.floor(y / zoom) * this.w;
for (let x = 0; x < W; x++) {
const c = rgb[this.buf[srow + Math.floor(x / zoom)]] || [255, 0, 255];
const o = (y * W + x) * 4;
d[o] = c[0]; d[o + 1] = c[1]; d[o + 2] = c[2]; d[o + 3] = 255;
}
}
return img;
}
}
export function hexToRgb(hex) {
const s = hex.replace('#', '');
return [parseInt(s.slice(0, 2), 16), parseInt(s.slice(2, 4), 16), parseInt(s.slice(4, 6), 16)];
}

163
js/selftest.js Normal file
View file

@ -0,0 +1,163 @@
// Assertions over the stages below detection. Runs in the browser so the exact
// module graph the tool uses is what gets tested.
//
// The ring-simplicity check exists because "fixed topology" is load-bearing in
// docs/roto-puppet.md: because hold parts CUT between poses rather than
// interpolating, a ring whose vertex order is wrong self-intersects and renders
// as blocks meeting at corners. It is invisible at some vertex counts and obvious
// at others, so it needs an assertion rather than an eyeball.
import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, subsampleSlots, subsampleRing } from './landmarks.js';
import { fitSimilarity, applySim, procrustesMean, smoothTransforms } from './mathutil.js';
import { stabilize, toRasterRing, selectKeys, activeKey } from './pipeline.js';
import { IndexedRaster, hexToRgb } from './raster.js';
import { writeTake } from './take.js';
import { synthDense } from './synth.js';
const results = [];
const ok = (name, cond, detail = '') => results.push({ name, pass: !!cond, detail });
/* ---- geometry helpers ---- */
function segmentsCross(a, b, c, d) {
const o = (p, q, r) => Math.sign((q.x - p.x) * (r.y - p.y) - (q.y - p.y) * (r.x - p.x));
const o1 = o(a, b, c), o2 = o(a, b, d), o3 = o(c, d, a), o4 = o(c, d, b);
return o1 !== o2 && o3 !== o4 && o1 !== 0 && o2 !== 0 && o3 !== 0 && o4 !== 0;
}
// A closed ring is simple if no pair of non-adjacent edges crosses.
function ringSelfIntersections(pts) {
const n = pts.length, hits = [];
for (let i = 0; i < n; i++) {
for (let j = i + 1; j < n; j++) {
if (i === j || (j + 1) % n === i || (i + 1) % n === j) continue;
if (segmentsCross(pts[i], pts[(i + 1) % n], pts[j], pts[(j + 1) % n])) hits.push([i, j]);
}
}
return hits;
}
const spreadX = (frames, slot) => {
const xs = frames.map((f) => f[slot].x);
return Math.max(...xs) - Math.min(...xs);
};
/* ---- the tests ---- */
export function run() {
results.length = 0;
// tables
ok('LIPS_OUTER has 20 distinct ids', new Set(LIPS_OUTER).size === 20);
ok('LIPS_INNER has 20 distinct ids', new Set(LIPS_INNER).size === 20);
ok('FACE_OVAL has 36 distinct ids', new Set(FACE_OVAL).size === 36);
ok('RIGID excludes every lip vertex',
!RIGID.some((i) => LIPS_OUTER.includes(i) || LIPS_INNER.includes(i)),
'a moving feature in the rigid set bleeds performance into stabilisation');
// subsampling preserves order and count at every budget
for (let n = 4; n <= 16; n += 2) {
const s = subsampleSlots(20, n);
const mono = s.every((v, i) => i === 0 || v > s[i - 1]);
ok(`subsampleSlots(20,${n}) is strictly increasing, n=${n}`, mono && s.length === n, s.join(','));
}
ok('subsampleRing agrees with subsampleSlots',
subsampleRing(LIPS_OUTER, 8).join(',') === subsampleSlots(20, 8).map((s) => LIPS_OUTER[s]).join(','));
const dense = synthDense(72);
// rings must be simple at EVERY vertex budget, on every frame
for (const [label, table] of [['outer', LIPS_OUTER], ['inner', LIPS_INNER]]) {
let worst = null;
for (let n = 4; n <= 16 && !worst; n += 2) {
const slots = subsampleSlots(table.length, n);
for (let f = 0; f < dense.length; f++) {
const pts = slots.map((s) => dense[f][table[s]]);
const hits = ringSelfIntersections(pts);
if (hits.length) { worst = `verts=${n} frame=${f} edges ${JSON.stringify(hits[0])}`; break; }
}
}
ok(`${label} ring is simple at every vertex budget`, !worst, worst || '');
}
// similarity fit recovers a known transform
const src = [{ x: 0, y: 0 }, { x: 1, y: 0 }, { x: 0, y: 1 }, { x: 2, y: 3 }];
const truth = { s: 1.7, theta: 0.6, tx: 4, ty: -2 };
const dst = src.map((p) => applySim(truth, p));
const got = fitSimilarity(src, dst);
ok('fitSimilarity recovers a known transform',
Math.abs(got.s - truth.s) < 1e-9 && Math.abs(got.theta - truth.theta) < 1e-9 &&
Math.abs(got.tx - truth.tx) < 1e-9 && Math.abs(got.ty - truth.ty) < 1e-9,
`s=${got.s.toFixed(6)} th=${got.theta.toFixed(6)}`);
// stabilisation: head motion out, mouth motion kept
const stab = stabilize(dense, 1);
const rawSpread = spreadX(dense, 133);
const stabSpread = (() => {
const xs = stab.eyes.map((e) => e[0].x);
return Math.max(...xs) - Math.min(...xs);
})();
ok('stabilisation removes >90% of head translation',
stabSpread < rawSpread * 0.1, `raw ${rawSpread.toFixed(4)} -> ${stabSpread.toFixed(4)}`);
const apRange = Math.max(...stab.aperture) - Math.min(...stab.aperture);
ok('stabilisation preserves mouth motion', apRange > 0.05, `aperture range ${apRange.toFixed(4)}`);
// key selection
const xf = (p) => ({ x: p.x * 320, y: p.y * 200 });
const shapes = stab.outer.map((r) => toRasterRing(r, LIPS_OUTER, 8, xf));
const sel = selectKeys(shapes, { minHold: 2, distThresh: 0.6, velSmooth: 3, exposure: 2 });
ok('keys are strictly increasing in f', sel.keys.every((k, i) => i === 0 || k.f > sel.keys[i - 1].f));
ok('keys respect the minimum hold',
sel.keys.every((k, i) => i === 0 || k.src - sel.keys[i - 1].src >= 2));
ok('keys land on the exposure grid', sel.keys.every((k) => k.f % 2 === 0));
ok('selection reduces candidates', sel.keys.length < sel.candidates.length,
`${sel.candidates.length} candidates -> ${sel.keys.length} keys`);
ok('first key is frame 0', sel.keys[0].f === 0);
ok('activeKey holds between keys',
activeKey(sel.keys, sel.keys[1].f - 1).f === sel.keys[0].f);
// rasteriser: indexed, hard-edged, no blending
const r = new IndexedRaster(64, 48);
r.clear(0);
r.fillPoly([{ x: 8, y: 8 }, { x: 56, y: 8 }, { x: 56, y: 40 }, { x: 8, y: 40 }], 2);
const present = new Set(r.buf);
ok('raster contains only written indices', present.size === 2 && present.has(0) && present.has(2),
`indices ${[...present].join(',')}`);
let count = 0;
for (const v of r.buf) if (v === 2) count++;
ok('axis-aligned rect fills the exact pixel count', count === 48 * 32, `${count} vs ${48 * 32}`);
const pal = ['#000000', '#ffffff', '#ff8800'];
const img = r.toImageData(pal, 2);
const seen = new Set();
for (let i = 0; i < img.data.length; i += 4) {
seen.add(`${img.data[i]},${img.data[i + 1]},${img.data[i + 2]}`);
}
const allowed = new Set(pal.map((h) => hexToRgb(h).join(',')));
ok('palette expansion introduces no intermediate colours',
[...seen].every((c) => allowed.has(c)), `${seen.size} distinct colours`);
// take writer round-trip
const take = {
name: 'test', frames: 72, width: 320, height: 200, exposure: 2,
palette: [{ name: 'bg' }, { name: 'skin' }],
slot: { x: 160, y: 100 },
parts: [
{ name: 'head', kind: 'plate', z: 0, interp: 'hold', keys: [{ f: 0, plate: 0 }] },
{ name: 'mouth', kind: 'poly', z: 30, color: 'skin', interp: 'hold',
keys: sel.keys.map((k) => ({ f: k.f, src: k.src, pts: shapes[k.src] })) },
],
};
const text = writeTake(take);
const keyLines = text.split('\n').filter((l) => l.startsWith('key') && l.includes('n='));
ok('every key line declares n= matching its point count',
keyLines.every((l) => {
const n = +l.match(/n=(\d+)/)[1];
const pts = l.split(/n=\d+\s+/)[1].trim().split(/\s+/);
return pts.length === n;
}), `${keyLines.length} key lines`);
ok('take declares a plate and a part table',
/^plate\s+0/m.test(text) && /^part\s+mouth/m.test(text));
ok('coordinates are integers', !/-?\d+\.\d/.test(text.split('\n').filter((l) => l.startsWith('key')).join('')));
return results;
}

69
js/synth.js Normal file
View file

@ -0,0 +1,69 @@
// Synthetic landmark frames, shaped exactly like FaceLandmarker output.
//
// Exists so the whole chain downstream of detection - Procrustes, smoothing,
// stabilisation, key selection, rasterising, take writing - can be exercised and
// verified without a video file. A synthetic face is also the only way to test
// stabilisation against a KNOWN head motion, since real footage gives no ground
// truth to compare against.
import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, EYE_INNER } from './landmarks.js';
const NUM = 478;
export function synthDense(nFrames = 72) {
const frames = [];
for (let t = 0; t < nFrames; t++) {
const pts = new Array(NUM);
for (let i = 0; i < NUM; i++) pts[i] = { x: 0.5, y: 0.5, z: 0 };
// Known head motion: drift, sway, roll and a slow scale change, plus a
// little per-frame jitter so transform smoothing has something to remove.
const ph = t / nFrames;
const hx = 0.5 + 0.045 * Math.sin(ph * Math.PI * 2) + (Math.random() - 0.5) * 0.002;
const hy = 0.5 + 0.02 * Math.cos(ph * Math.PI * 3) + (Math.random() - 0.5) * 0.002;
const roll = 0.18 * Math.sin(ph * Math.PI * 2.5);
const scale = 1 + 0.06 * Math.sin(ph * Math.PI * 1.5);
const cr = Math.cos(roll), sr = Math.sin(roll);
const place = (i, lx, ly) => {
const sx = lx * scale, sy = ly * scale;
pts[i] = { x: hx + cr * sx - sr * sy, y: hy + sr * sx + cr * sy, z: 0 };
};
// Mouth opens in four sustained beats with holds between, so key selection
// has genuine extremes and genuine plateaux to find.
const beat = Math.floor(t / 9) % 4;
const target = [0.004, 0.05, 0.022, 0.0];
const openAmt = target[beat];
const wide = 0.10 + (beat === 1 ? 0.012 : beat === 3 ? -0.008 : 0);
place(RIGID[0], -0.075, -0.045); place(RIGID[1], -0.028, -0.043);
place(RIGID[2], 0.028, -0.043); place(RIGID[3], 0.075, -0.045);
place(RIGID[4], 0.000, -0.050); place(RIGID[5], 0.000, -0.020);
place(RIGID[6], 0.000, 0.012);
place(EYE_INNER[0], -0.028, -0.043); place(EYE_INNER[1], 0.028, -0.043);
// Lip rings as ellipse arcs, traversed so ring ORDER matches the tables:
// slot 0 = right corner, 5 = top centre, 10 = left corner, 15 = bottom
// centre, with y growing downward. Getting this convention wrong swaps two
// opposite vertices and the ring self-intersects into a bowtie - see the
// ring-simplicity assertion in selftest.
const ring = (table, rx, ry, cy) => {
const n = table.length;
for (let k = 0; k < n; k++) {
const a = -(k / n) * Math.PI * 2;
place(table[k], rx * Math.cos(a), cy + ry * Math.sin(a));
}
};
ring(LIPS_OUTER, wide / 2, 0.012 + openAmt * 0.6, 0.075);
// APERTURE (13, 14) are slots 5 and 15 of the inner ring, so the ring itself
// places them at the vertical extremes. Writing them again afterwards is what
// produced the bowtie; the aperture is simply the inner ring's height.
ring(LIPS_INNER, wide / 2.6, 0.001 + openAmt, 0.075);
for (let k = 0; k < FACE_OVAL.length; k++) {
const a = -Math.PI / 2 + (k / FACE_OVAL.length) * Math.PI * 2;
place(FACE_OVAL[k], 0.105 * Math.cos(a), 0.145 * Math.sin(a) + 0.01);
}
frames.push(pts);
}
return frames;
}

41
js/take.js Normal file
View file

@ -0,0 +1,41 @@
// Take-file writer. Format is specified in docs/roto-puppet.md, "The take
// format". Line-oriented text on purpose: Poco can parse it with fopen/fgets
// from poco/src/safefile.c and strtok/atoi/atof from poco/src/strlib.c, so the
// Animator Pro render script needs no new native code.
const r = (v) => Math.round(v);
export function writeTake(take) {
const L = [];
L.push(`take name=${take.name} frames=${take.frames} width=${take.width} height=${take.height} exposure=${take.exposure}`);
L.push('');
take.palette.forEach((p, i) => L.push(`pal ${i} ${p.name.padEnd(11)} ${i}`));
L.push('');
// v1 emits a single frozen plate derived from the face oval. A real project
// replaces this with hand-drawn angles referenced by cel frame; the record
// shape is the same either way.
L.push(`plate 0 kind=poly slot_mouth=${r(take.slot.x)},${r(take.slot.y)} scale=1.00 rot=0 squash=1.00`);
L.push('');
for (const part of take.parts) {
const bits = [`part ${part.name.padEnd(9)} kind=${part.kind} z=${part.z}`];
if (part.color !== undefined) bits.push(`color=${part.color}`);
if (part.kind === 'poly') bits.push('closed=1 fill=1');
bits.push(`interp=${part.interp}`);
if (part.parent) bits.push(`parent=${part.parent}`);
L.push(bits.join(' '));
}
L.push('');
for (const part of take.parts) {
for (const k of part.keys) {
if (k.hidden) { L.push(`key ${part.name.padEnd(9)} f=${k.f} hidden`); continue; }
if (part.kind === 'plate') { L.push(`key ${part.name.padEnd(9)} f=${k.f} plate=${k.plate}`); continue; }
const pts = k.pts.map((p) => `${r(p.x)},${r(p.y)}`).join(' ');
// f is authoritative (what renders); src is the pre-snap extreme frame,
// kept only as a tuning signal - a key dragged far means the minimum-hold
// and the exposure grid are fighting.
L.push(`key ${part.name.padEnd(9)} f=${k.f} src=${k.src} n=${k.pts.length} ${pts}`);
}
}
L.push('');
return L.join('\n');
}

20
selftest.html Normal file
View file

@ -0,0 +1,20 @@
<!doctype html>
<html lang="en"><head><meta charset="utf-8"><title>selftest</title>
<style>
body{background:#0b0d13;color:#e8eaf0;font:13px/1.6 ui-monospace,Menlo,monospace;padding:20px}
li{list-style:none} .p{color:#4ade80} .f{color:#f87171}
b{font-weight:400;color:#8891a5} h1{font-size:14px}
</style></head><body>
<h1 id="head">running…</h1><ul id="out"></ul>
<script type="module">
import { run } from './js/selftest.js';
let res;
try { res = run(); }
catch (e) { res = [{ name: 'harness threw: ' + e.message, pass: false, detail: String(e.stack).split('\n')[1] || '' }]; }
const pass = res.filter(r => r.pass).length;
const fail = res.length - pass;
document.getElementById('head').textContent = `${fail === 0 ? 'PASS' : 'FAIL'} — ${pass}/${res.length} assertions`;
document.title = `${fail === 0 ? 'PASS' : 'FAIL'} ${pass}/${res.length}`;
document.getElementById('out').innerHTML = res.map(r =>
`<li class="${r.pass ? 'p' : 'f'}">${r.pass ? '✓' : '✗'} ${r.name}${r.detail ? ` <b>${r.detail}</b>` : ''}</li>`).join('');
</script></body></html>