Compare commits

...

10 commits

Author SHA1 Message Date
Your Name
c2603102da readme additions 2026-09-24 15:49:05 -04:00
f827757f7d Become arthur: a standalone suite, not an Animator Pro front-end
The test renderer turned out to be the product. Everything that decides how the
work looks - stabilisation, reduction, timing, frame removal, palette - already
happens here, and the flat indexed output already reads the way it should.

The reason to leave is in the original design's own rule: never make a timing
decision that requires a full render to evaluate. Honouring that moved every
judgement out of Animator Pro, which left the host doing nothing but writing a
file, in exchange for modal UI, minutes-long renders, one-level undo, FLX delta
invariants, a single tween state and a cel singleton.

What does NOT change is the constraint. 320x200, indexed palette, flat fills,
no antialiasing - inherited, but load-bearing rather than accidental. The
rasteriser writes palette indices and expands to RGBA only at the end precisely
so nothing can soften an edge. Modern conveniences belong in the workflow.

Adds docs/design.md: the principles, carried over without the Poco/FLX/cel
machinery, plus architecture and an honest list of what is missing - the
largest gap being that plates still have nowhere to be drawn.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:47:41 -04:00
dc0a3f34e6 Make the mouth lead legible and assertable
"It feels like it does nothing" is indistinguishable from "it does nothing",
and at 24fps a lead of 1 is 42ms - small enough to reasonably doubt. So the
shift is now provable and visible rather than taken on trust:

- shiftIndex is a pure exported function with assertions covering identity,
  both directions, and clamping at each end
- the frame label always shows the mouth frame, not only when shifted, so the
  number can be watched diverging from f; non-zero leads also report in ms
- the slider readout carries an explicit + sign
- the stabilised pane draws the unshifted contour as a dark-green ghost when a
  lead is set, so the offset is something you can see

Also: leadIndex called opts() on every invocation - ~14 DOM reads, once per
strip thumbnail, so ~1000 per redraw at 74 frames. It reads a cached scalar now.

The "different pose" assertion checks the whole track rather than one pair:
synthetic poses hold for nine-frame beats, so a single pair can legitimately be
identical while the shift works correctly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:38:02 -04:00
f9b8ec8617 Fix teeth vertex slider driving the lip vertex count
opts() declared `verts` twice - once from the lip slider and again from the
teeth slider. Duplicate keys in an object literal are silent in JS and the last
one wins, so the teeth vertex control was quietly setting the lip vertex budget
while the lip control did nothing at all.

Renamed to teethVerts, with interior.js reading it under that name.

selftest now parses the opts() literal and fails on duplicate keys, verified to
catch this exact case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:36:04 -04:00
d18791e6ad Mouth lead: nudge performance tracks ahead of the clock
Smoothing does not literally delay anything - a centred moving average has zero
phase lag - but it blurs onsets, so the visually salient moment of a mouth
opening moves later even though the mean does not. Animators also draw mouth
shapes a frame or two ahead of the sound as standard practice, so this is the
normal control rather than a workaround.

Only the performance tracks shift; the head and audio stay put, since it is the
mouth that should anticipate. In photo-underlay modes the vector mouth will
therefore no longer match the frame behind it, which is expected.

The lead is baked into the exported take - key f carries the pose from source
frame f+lead - so the Animator Pro renderer never needs to know about it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:34:08 -04:00
a146d271b6 Cross-check el() ids against index.html in the selftest
A knob wired in app.js but missing from the markup throws during wiring, aborts
the module and leaves a blank page - a symptom pointing nowhere near its cause.
It has now happened twice, so it gets a check rather than vigilance: the
selftest fetches both files and compares the id sets in each direction.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:25:39 -04:00
9a11eeabc1 Teeth as an extracted blob contour, not a clipped band
The band filled the mouth because a band is the wrong reduction: the bright
region is a blob, and reading it as "everything above a line" throws the shape
away.

Extracting a contour reintroduces the vertex-correspondence problem that made
me avoid it, but for a blob there is a way out. Radial sampling from the
centroid along N fixed directions makes vertex k always mean "the extent in
direction k": correspondence holds by construction, the count is fixed, and
temporal smoothing cannot reorder anything. It also yields a star-shaped
reduction, which suits flat colour.

Tongue rejection, which the band had no way to express:
- pixels red relative to their own brightness are dropped (teeth are neutral)
- component choice is biased toward the top of the cavity, since area alone
  picks the tongue when the mouth is wide
- separate inner and outer controls: cavity erode pulls the sampled region off
  the lip edge, blob grow/erode resizes the found blob

Also: a knob wired in app.js but missing from index.html threw during wiring
and left a blank page with nothing useful in the console - which is exactly
what happened to teethDwell in the previous commit. el() now names the missing
id, and window.onerror surfaces it in the status line.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:24:52 -04:00
c6daf827d1 Fix teeth band filling the whole mouth
Three causes, all of them mine:

Otsu always returns a split, including on a homogeneous region - given a dark
cavity with no teeth it invents a threshold and calls half the pixels bright.
The gate is now the separation between the two class means, which is the only
thing that says whether the split means anything. Coverage was the wrong
signal: it is high both when the mouth is full of teeth and when the region is
uniformly dark and Otsu has split noise.

The row scan tracked the last qualifying row anywhere rather than where the run
from the top stops, so one bright row near the bottom - a lit lower lip inside
the ring - pushed the line to full height. It now breaks at the first failing
row once the run has started.

MediaPipe's inner lip landmarks sit slightly outside the real opening, so the
sampled region included lip pixels, which are bright and sit exactly at the
boundary where they do most damage. The ring is now eroded toward its centroid
before sampling, with the amount exposed as a knob.

Adds a diagnostic panel showing the sampled crop, pixels above threshold, and
the resolved line, because tuning this from numbers alone does not tell you
whether the region being measured is even the right region.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:14:49 -04:00
941022b69f Teeth from mouth-interior image content
MediaPipe has no landmarks inside the lips, so teeth have to come from pixels.
Tracing the bright blob would give a new contour every frame with no vertex
correspondence - the exact boil docs/roto-puppet.md warns about.

So the measurement yields a scalar, not a shape: Otsu within the cavity, scanned
from the top for where the bright run stops, giving one line height per frame.
The teeth polygon is the inner lip ring clipped to that line, so the silhouette
is always the mouth's own shape and cannot disagree with the lips around it,
and the only per-frame variable is a single number that smooths trivially.

Presence gets hysteresis and minimum dwell, as plate selection does: a teeth
block blinking on and off for single frames is worse than one simply absent.

Tongue is not implemented. The same scalar approach would apply, gated on
redness rather than brightness, but it is not visible in the test footage - the
cavity reads dark with a bright upper-teeth band and nothing else.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:12:09 -04:00
29b0bd6690 Fix horizontal stretch from MediaPipe's anisotropic normalised space
MediaPipe normalises x by image WIDTH and y by image HEIGHT, so for a 1080x1920
clip one unit of x is 1080px and one unit of y is 1920px. makeXform applied a
single scale to both, stretching everything horizontally by H/W - 1.78x on this
footage. The photo underlay looked equally squashed because frameAffine divided
x by imgW, matching the equally wrong vector shapes rather than disagreeing
with them.

Fixed at ingest: landmarks convert to an isotropic space whose unit is one image
height (x *= W/H), so equal numbers mean equal pixels everywhere downstream.
Pixel mapping follows - both axes divide by imgH.

This also silently fixes head roll. fitSimilarity was fitting a rotation in a
sheared space, so the "similarity" it recovered was not one, and stabilisation
of rolled heads was subtly wrong.

selftest: a shape circular in pixel space must stay circular in raster space,
checked at 1080x1920, 1920x1080 and 640x640. Fails at ratio 1.78 without the
conversion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:03:12 -04:00
14 changed files with 921 additions and 137 deletions

View file

@ -1,16 +1,18 @@
# roto
# arthur
Video → **take file** builder for the Animator Pro rotoscope/puppet pipeline
described in `../docs/roto-puppet.md`.
An animation suite for turning live-action video into work that reads as
hand-authored — flat shapes, a tiny palette, hard edges, motion carried by
silhouette, in the idiom of *Another World*.
This is the analysis and tuning half. It stabilises a face out of a clip, reduces
the lip contour to a handful of vertices, selects sparse keys on motion extremes,
and previews the result as flat indexed fills — so the *look and timing* can be
judged in seconds rather than through a minutes-long Animator Pro render. It
emits a `.take` file; nothing here touches Animator Pro.
It stabilises a face out of a clip, reduces the lip contour to a handful of
vertices, derives teeth from image content, lets you decide which frames need
their own hand-drawn head, and renders the result as flat indexed fills at
320×200 with no antialiasing.
All policy lives here. The take arrives at the renderer with its keys already
chosen.
The constraint is deliberate. 320×200 and an indexed palette are inherited from
Animator Pro, where this started, but they are why the output looks right —
modern conveniences belong in the workflow, not the output. See
[docs/design.md](docs/design.md).
## Run
@ -19,6 +21,9 @@ python3 -m http.server 8777 # from this directory
# open http://127.0.0.1:8777
```
Static files and ES modules — no build step, no dependencies beyond MediaPipe's
wasm, which is fetched from a CDN on first use.
**Synthetic take** needs no video and exercises everything below detection.
For real footage:
@ -57,11 +62,43 @@ is hand-drawn head plates, which this tool does not yet do.
| Knob | What it does |
| --- | --- |
| vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. |
| mouth lead ±f | Shifts the performance tracks earlier against the audio and the head. `[` `]`. |
| contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. |
| anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform*. |
| closed-mouth cut | Aperture below which the mouth interior is emitted as `hidden`. |
| suggest tolerance | Max head movement before a new plate drawing is required. Affects **Suggest** only. |
## Teeth
MediaPipe has no landmarks inside the lips: the inner ring bounds the cavity and
everything within it is just pixels. So teeth come from the image.
The hazard is vertex correspondence — a traced contour reorders between frames
and boils. The way out for a blob specifically is **radial sampling**: march
outward from the blob's centroid along N fixed directions and take the last pixel
inside. Vertex *k* is then always "the extent in direction *k*", so correspondence
holds by construction, the vertex count is fixed, and temporal smoothing is well
defined with no reordering possible. It also produces a star-shaped reduction,
which is what flat blocks of colour want.
The **teeth measurement** panel shows exactly what is sampled: region dimmed,
kept pixels green, extracted contour amber. Tune against that, not the numbers.
| Knob | What it does |
| --- | --- |
| teeth contrast | Gate on the separation between the cavity's dark and bright class means. Otsu always returns *some* threshold, so this is what stops it inventing teeth in a dark mouth. |
| cavity erode | Pulls the sampled region in from the lip edge — MediaPipe's inner lip landmarks sit slightly outside the real opening, and lips are bright. |
| blob grow/erode | Resizes the found blob. An open pass always runs first to despeckle. |
| tongue reject | Drops pixels red relative to their own brightness. Teeth are near-neutral; tongue is not. |
| prefer upper | Biases component choice toward the top of the cavity. Area alone picks the tongue when the mouth is wide. |
| teeth vertices | Radial sample count. |
| teeth avg ±f | Temporal average over the contour. |
| teeth dwell | Frames a presence change must persist before it takes effect. |
**Tongue** as its own part would work the same way, gated on redness instead of
brightness and biased low rather than high. Not implemented: it is not visible in
the test footage, which reads as a dark cavity with a bright upper-teeth band.
## The plate is reference, not art
The plate layer has several representations because its job changes. Cycle with
@ -103,8 +140,16 @@ their own plate drawing. Everything starts kept; delete what you don't want.
**Suggest** runs error-tolerance decimation over head pose as a starting point,
then you hand-correct.
**Mouth lead** is not a correction for a bug. A centred moving average has no
phase lag, so smoothing does not delay anything — but it blurs onsets, and the
visually salient moment of a mouth opening moves later even though the mean does
not. Animators also draw mouth shapes one or two frames ahead of the sound as
standard practice. Only the performance tracks shift; the head stays with the
audio, since it is the mouth that should anticipate. The lead is baked into the
exported take, so the renderer never needs to know about it.
`contour avg` is a deliberate, bounded exception to "never smooth the contour" in
`../docs/roto-puppet.md`. That rule held while keys were sparse, because sampling
`docs/design.md`. That rule held while keys were sparse, because sampling
at velocity minima rejected detector noise for free. With a key on every frame it
does not, so a radius shorter than the shortest articulation worth keeping is
justified — at 12fps, articulation spans 3–6 frames and detector noise is
@ -117,7 +162,11 @@ chromium --headless --virtual-time-budget=8000 --dump-dom \
http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+'
```
Or open `selftest.html`. 29 assertions over the stages below detection.
Or open `selftest.html`. 41 assertions over the stages below detection, plus a
wiring cross-check: every `el('id')` in `app.js` must exist in `index.html`. A
knob wired in one but not the other throws during wiring, which aborts the rest
of the module and leaves a blank page — a symptom that points nowhere near its
cause, and which has happened twice.
The ring-simplicity check is the load-bearing one. Because `hold` parts *cut*
between poses instead of interpolating, a ring whose vertex order is wrong

219
docs/design.md Normal file
View file

@ -0,0 +1,219 @@
# arthur — design
A small animation suite for turning live-action video into work that reads as
hand-authored, in the idiom of *Another World* (Éric Chahi, 1991): flat shapes,
a tiny palette, hard edges, motion carried by silhouette.
## The aesthetic is a representation, not a filter
Chahi did not process video. He shot reference footage and hand-traced polygons
over it in a custom editor; the engine stored and replayed polygon lists, never
bitmaps. Three properties follow, and all three are load-bearing:
1. **Temporal identity.** The same shape, with the same vertex count and vertex
order, *edited* across frames. A shape re-detected independently each frame
produces a new contour every frame. That boils, and boiling reads as "filter"
within half a second no matter how good the individual shapes are.
2. **Stylised timing.** Poses held, and for most parts a hard cut between them
rather than an interpolation. A distinct pose on every frame of 24fps footage
looks like video even when every pose is a polygon.
3. **Authored colour.** A small fixed ramp with two or three tones per part,
chosen by a person. Sampling colour from the source produces a pixel-art
filter immediately and irrecoverably.
None of those are computer-vision problems. That is the whole argument about
where CV belongs.
## The constraint is the point
320×200, indexed palette, flat fills, no antialiasing. That is inherited from
Animator Pro, where this work started, but it is not an accident to be
modernised away — it is why the output looks right. The rasteriser writes palette
indices into a byte buffer and expands to RGBA only at the end, precisely so no
canvas antialiasing can soften an edge.
Modern conveniences belong in the *workflow* — instant feedback, audio, real
undo, scrubbing, layers. Not in the output.
## CV tracks; it does not draw
- **Its job.** Where the rigid features of the face are, what the head's
transform is, where the lip contour runs, which frame is a motion extreme.
- **Its non-job.** Deciding what the shapes are. A face mesh offers 468 points;
an animator's mouth is eight. **That decimation ratio is the style.** It is a
decision encoded in the tool, not a measurement extracted from footage.
Every number crossing from analysis into rendering is a *parameter* or a
*correspondence*. Every number determining how something looks comes from the
artist or a fixed authored table.
## Two kinds of part
| Kind | Source | Vocabulary | Interp |
| --- | --- | --- | --- |
| **Plate** — head, hair, body | Hand-drawn | Closed: a few drawings per character | hold |
| **Feature** — mouth, lids | Rotoscoped from landmarks | Open: derived from this take | hold |
| **Interior** — mouth interior, teeth | Image content within a feature | Open | hold |
| **Primitive** — iris | Landmark centroid as a disc | Quantised | hold |
The asymmetry is deliberate, and it is the opposite choice in each case.
A **closed vocabulary is wrong for the mouth.** Pre-authored mouth shapes are
never *this performance's* shapes, and that genericness is what the project
exists to avoid. So the mouth gets an open vocabulary derived from footage, and
the stylisation applies to its **timing** instead.
A **closed vocabulary is right for the head.** The head carries structure, not
performance. A handful of hand-drawn angles is better, because you drew them —
and it means the head never needs segmentation, contour tracking, or boil
avoidance.
## Two kinds of sparseness
Conflating these was the original design error. **Aesthetic** sparseness is set
by the extraction rate: pick 12fps and the timing is already chosen. **Labour**
sparseness is a human drawing each one, and it binds only on the plate.
So the mouth keeps **every** frame — it is traced, and therefore free. In limited
animation lip sync is routinely the densest element, on 1s, while heads hold on
2s and 3s.
The other half is frame removal: which frames need their own plate drawing.
Everything starts kept; delete what you don't want. Automatic suggestion runs
error-tolerance decimation over head pose as a starting point, then you
hand-correct.
## Stabilisation
A rotoscoped mouth is only reusable if expressed independently of where the head
was. That inverts the obvious composition:
```text
mouth_local(t) = anchor(t)⁻¹ · mouth_world(t)
```
- **Rigid landmarks only** for the fit: eye corners, nose bridge, nose tip.
Including a feature that moves bleeds performance into the stabilisation.
- **Similarity, not affine or homography.** The extra degrees of freedom absorb
head rotation as distortion and smear it into the mouth. Four DOF removes
exactly translation, roll and depth scale, leaving yaw and pitch as a
measurable residual.
- **Reference is the mean** configuration over the shot, not frame zero.
- **Smooth the transform, never the contour** — with one bounded exception,
below.
- Landmarks convert to an **isotropic** space first (unit = one image height).
MediaPipe normalises x by width and y by height, so its space is stretched; a
"similarity" fitted there is not one.
Out-of-plane rotation cannot be removed by any 2D transform. The answer is not a
better transform: it is to draw the head at each angle and let each drawing
declare where its mouth sits. Foreshortening becomes authored metadata.
## Hold versus interpolate
Another World had no inbetweening engine: polygon sets played back frame by
frame at a low rate. For parts that should read as hand-animated snaps — mouths
above all — **cut between keys, do not interpolate.** A tweened mouth is rubbery
and reads as puppet software immediately.
**Fixed topology is non-negotiable for cut parts.** Vertex *meanings* must be
stable across every key, so a closed mouth and a wide-open mouth are the same
polygon at different positions. Interpolation partially hides ordering drift;
cutting exposes it completely, as static.
## Deriving shapes from pixels without boiling
Some things have no landmarks — teeth, tongue, anything inside the lips. The
hazard is vertex correspondence: a traced contour reorders between frames.
Two escapes, both used:
- **Extract a scalar, not a shape**, where the shape can be derived from
geometry you already trust.
- **Radial sampling** where a real contour is needed. March outward from the
blob's centroid along N fixed directions; vertex *k* is then always "the extent
in direction *k*". Correspondence holds by construction, the count is fixed,
temporal smoothing is well defined, and the star-shaped result suits flat
colour.
## The bounded smoothing exception
*Smooth the transform, never the contour* held while keys were sparse: sampling
at velocity minima rejected per-frame detector noise for free. With a key on
every frame it does not, so a short contour average is justified — the window
must stay **shorter than the shortest articulation worth keeping**. At 12fps,
mouth movement spans 3–6 frames and detector noise is per-frame, so ±1 separates
them and ±3 starts eating speech. At 24fps, double it.
## Performance tracks may lead the clock
A centred moving average has no phase lag, but it blurs onsets, so the visually
salient moment of a mouth opening moves later even though the mean does not.
Animators also draw mouth shapes a frame or two ahead of the sound as standard
practice. Only performance tracks shift; the head stays with the audio, because
it is the mouth that should anticipate.
## Colour discipline
- A small ramp. Another World ran 16 colours at 320×200.
- Two or three tones per part: base, shadow, occasionally a rim.
- No dithering, no antialiasing anywhere.
- Never sample colour from the source. Analysis may report *which* tone a region
should be; it must never report an RGB value.
- Plate art and generated parts share one palette, authored together.
## Three artifacts, not two
| Artifact | Cost | Regenerated when |
| --- | --- | --- |
| Dense track — landmarks and fitted anchor per frame | Minutes, once per shot | Re-shoot, or a model change |
| Take — parts, keys, kept frames | Instant | Every knob change |
| Render | — | Continuous |
Keep the dense track: re-keying is then a regenerate, never a re-trace. And
because the take is disposable, **hand corrections must not live in it** — they
belong in an override layer keyed by `(part, frame)`, applied on top at render
time.
## Architecture
The organising idea is an **indexed raster core**, a **part/track model**, and
**pluggable sources** feeding it. Sources so far are rotoscope and image-derived
interior; hand-drawn and procedural are the obvious next ones. Everything else —
stabilisation, key selection, frame removal, palette — operates on the track
model and is source-agnostic.
Current modules:
| Module | Role |
| --- | --- |
| `landmarks.js` | Index tables. Ring arrays are ordered traversals: slot position *is* vertex identity. |
| `mathutil.js` | Similarity fit, Procrustes mean, temporal smoothing. |
| `pipeline.js` | Stabilise → subsample → key-select → frame-removal. |
| `interior.js` | Teeth from image content: Otsu, morphology, components, radial contour. |
| `underlay.js` | Registered photo reference and palette posterisation. |
| `raster.js` | Indexed scanline fill. No antialiasing, by construction. |
| `take.js` | Take-file writer. |
| `synth.js` | Synthetic landmarks, so everything below detection is testable without a video. |
| `selftest.js` | Assertions, including source-level wiring checks. |
## Not built yet
- **A paint surface.** The plates have nowhere to be drawn. This is the largest
gap between "tool" and "suite": a pixel paint canvas with onion skin, palette
constraint, and the registered underlay behind it.
- Eyes, irises, brows as parts.
- Plate libraries with per-plate mouth slots.
- Real performer→character calibration (currently identity).
- The override layer.
- Export beyond the take file — FLI/FLC would connect back to the lineage and is
a simple format.
## Lineage
This began as an analysis front-end for Animator Pro, whose render path was the
intended destination. `Animator-Pro-SDL/docs/roto-puppet.md` documents that
design and what the Poco/FLX machinery can and cannot do. The reason for leaving
is in that document's own rule — *never make a timing decision that requires a
full render to evaluate* — which, once honoured, left the host with nothing to do
but write the file.

View file

@ -1,6 +0,0 @@
<!doctype html><html><body><pre id=o>?</pre><script>
const c=document.createElement('canvas');
const g=c.getContext('webgl2')||c.getContext('webgl');
document.getElementById('o').textContent = g ? 'WEBGL OK '+g.getParameter(g.VERSION) : 'WEBGL UNAVAILABLE';
document.title = g ? 'GLOK' : 'GLNO';
</script></body></html>

View file

@ -3,7 +3,7 @@
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>roto — take builder</title>
<title>arthur</title>
<style>
:root { --bg:#0b0d13; --panel:#12151f; --line:#232836; --fg:#e8eaf0;
--dim:#8891a5; --accent:#fbbf24; --ok:#4ade80; --err:#f87171; }
@ -58,7 +58,7 @@
</head>
<body>
<header>
<h1>roto <span>— frame removal + take builder</span></h1>
<h1>arthur <span>— rotoscope + frame removal</span></h1>
<button id="btn-synth">Synthetic</button>
<input type="text" id="framedir" value="frames" size="7" title="frame directory">
<button id="btn-frames">Load frames</button>
@ -91,7 +91,7 @@
<div class="panel">
<h2>stabilised (head-local)</h2>
<canvas id="cv-stab"></canvas>
<div class="legend">should sit still except the mouth · grey = held plate outline</div>
<div class="legend">should sit still except the mouth · grey = held plate outline<br>dark green ghost = unshifted mouth when lead ≠ 0</div>
</div>
<div class="panel">
<h2>flat render — 320×200 indexed</h2>
@ -109,6 +109,7 @@
<input type="range" id="scrub" min="0" max="0" value="0" style="width:100%;margin-top:10px">
<div class="legend">
<kbd>←</kbd> <kbd>→</kbd> step · <kbd>X</kbd> delete · <kbd>K</kbd> keep · <kbd>B</kbd> background ·
<kbd>[</kbd> <kbd>]</kbd> lead ·
click to select · double-click or shift-click to toggle ·
the mouth keeps <b>every</b> frame regardless
</div>
@ -119,19 +120,47 @@
<div class="panel" style="flex:1 1 400px">
<h2>knobs</h2>
<label class="ctl"><span>vertices</span><input type="range" id="verts" min="4" max="16" step="2" value="8"><output id="vertsv"></output></label>
<label class="ctl"><span>mouth lead ±f</span><input type="range" id="lead" min="-6" max="6" value="0"><output id="leadv"></output></label>
<label class="ctl"><span>contour avg ±f</span><input type="range" id="contourSmooth" min="0" max="4" value="1"><output id="contourSmoothv"></output></label>
<label class="ctl"><span>anchor avg ±f</span><input type="range" id="smoothWin" min="0" max="8" value="2"><output id="smoothWinv"></output></label>
<label class="ctl"><span>closed-mouth cut</span><input type="range" id="apertureThresh" min="0" max="400" value="120"><output id="apertureThreshv"></output></label>
<label class="ctl"><span>teeth contrast</span><input type="range" id="teethOn" min="1" max="60" value="16"><output id="teethOnv"></output></label>
<label class="ctl"><span>cavity erode</span><input type="range" id="teethErode" min="0" max="45" value="18"><output id="teethErodev"></output></label>
<label class="ctl"><span>blob grow/erode</span><input type="range" id="blobGrow" min="-4" max="4" value="0"><output id="blobGrowv"></output></label>
<label class="ctl"><span>tongue reject</span><input type="range" id="tongueReject" min="2" max="40" value="18"><output id="tongueRejectv"></output></label>
<label class="ctl"><span>prefer upper</span><input type="range" id="topBias" min="0" max="120" value="60"><output id="topBiasv"></output></label>
<label class="ctl"><span>teeth vertices</span><input type="range" id="teethVerts" min="5" max="20" value="10"><output id="teethVertsv"></output></label>
<label class="ctl"><span>teeth avg ±f</span><input type="range" id="teethSmooth" min="0" max="4" value="1"><output id="teethSmoothv"></output></label>
<label class="ctl"><span>teeth dwell</span><input type="range" id="teethDwell" min="0" max="6" value="1"><output id="teethDwellv"></output></label>
<label class="ctl"><span>suggest tolerance</span><input type="range" id="tol" min="2" max="60" value="14"><output id="tolv"></output></label>
<div class="legend">
<b>mouth lead</b> shifts the performance tracks earlier (positive) against
the audio and the head. Averaging has no phase lag but it blurs onsets, so
an opening reads later than it is; animators also draw mouths a frame or
two ahead of the sound as standard practice. <kbd>[</kbd> <kbd>]</kbd>.<br>
<b>contour avg</b> 0 = off, 1 = ±1 frame. Removes per-frame landmark
jitter. Push past 2 and it starts eating articulation.<br>
<b>anchor avg</b> smooths the head transform only — never the contour.<br>
<b>teeth contrast</b> gates on how far apart the cavity's dark and bright
halves are — Otsu always returns <i>some</i> threshold, so this is what
stops it inventing teeth in a dark mouth.
<b>cavity erode</b> pulls the sampled region in from the lip edge;
<b>blob grow/erode</b> resizes the found blob itself.
<b>tongue reject</b> drops pixels that are red relative to their own
brightness; <b>prefer upper</b> biases component choice toward the top of
the cavity, where teeth are and the tongue is not.
<b>dwell</b> is how many frames a presence change must persist.<br>
<b>suggest tolerance</b> only affects the Suggest button: max head movement
allowed before a new drawing is required.
</div>
</div>
<div class="panel" style="flex:1 1 280px">
<div class="panel" style="flex:0 1 190px">
<h2>teeth measurement</h2>
<div id="cv-teeth"></div>
<div class="legend" id="teethinfo"></div>
<div class="legend">green = kept pixels · amber = extracted contour</div>
</div>
<div class="panel" style="flex:1 1 240px">
<h2>palette</h2>
<div id="palette"></div>
<div class="legend" style="margin-top:12px">

217
js/app.js
View file

@ -1,8 +1,10 @@
import { FaceLandmarker, FilesetResolver } from 'https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@1.0.1/vision_bundle.mjs';
import { LIPS_OUTER, LIPS_INNER, FACE_OVAL } from './landmarks.js';
import { stabilize, toRasterRing, smoothContours, suggestPlateFrames, heldFrame } from './pipeline.js';
import { stabilize, toRasterRing, smoothContours, suggestPlateFrames, heldFrame, shiftIndex } from './pipeline.js';
import { IndexedRaster } from './raster.js';
import { drawRegistered, posterizeInto } from './underlay.js';
import { extractTeeth } from './interior.js';
import { applySim } from './mathutil.js';
import { writeTake } from './take.js';
import { synthDense } from './synth.js';
@ -13,9 +15,9 @@ const PALETTE = [
{ name: 'skin_base', hex: '#b07a5a' },
{ name: 'skin_dark', hex: '#7a4f3a' },
{ name: 'mouth_dark', hex: '#24161a' },
{ name: 'skin_lite', hex: '#d9a884' },
{ name: 'teeth', hex: '#d9cfc2' },
];
const IDX = { bg: 0, base: 1, dark: 2, mouth: 3, lite: 4 };
const IDX = { bg: 0, base: 1, dark: 2, mouth: 3, teeth: 4 };
const state = {
dense: null, images: [], stab: null, xform: null,
@ -23,11 +25,30 @@ const state = {
keep: new Set(), // frames that get their own plate drawing
frame: 0, playing: false, faceBox: null,
fps: 12, audio: null, // fps comes from manifest.json, never guessed
aspect: 1, // imgW/imgH; converts MediaPipe's anisotropic space
lead: 0, // performance-track offset in frames
interior: null, // per-frame teeth measurement from image content
teeth: null, // resolved per-frame {show, t} after knobs
};
const el = (id) => document.getElementById(id);
const el = (id) => {
const n = document.getElementById(id);
// A knob present in the code but missing from the markup used to throw during
// wiring and leave a blank page with nothing in the console worth reading.
if (!n) throw new Error(`missing element #${id} — knob wired in app.js but not in index.html`);
return n;
};
const opts = () => ({
verts: +el('verts').value,
lead: +el('lead').value,
teethOn: +el('teethOn').value / 100, // minimum Otsu class separation
teethDwell: +el('teethDwell').value,
teethSmooth: +el('teethSmooth').value,
cavityErode: +el('teethErode').value / 100,
tongueReject: +el('tongueReject').value / 100,
blobGrow: +el('blobGrow').value,
topBias: +el('topBias').value / 100,
teethVerts: +el('teethVerts').value,
smoothWin: +el('smoothWin').value,
contourSmooth: +el('contourSmooth').value,
apertureThresh: +el('apertureThresh').value / 1000,
@ -115,6 +136,19 @@ async function detectAll(images) {
return { dense, missing };
}
// Interior measurement is a function of pixels alone, so it runs once with
// detection and the knobs re-resolve it instantly afterwards.
function measureAll(images, dense, o) {
const ctx = document.createElement('canvas').getContext('2d', { willReadFrequently: true });
return dense.map((lm, i) =>
extractTeeth(images[i], LIPS_INNER.map((k) => lm[k]), ctx, o));
}
// Extraction keys on every knob that changes the pixels examined, so the cache
// is keyed on exactly those and a change to anything else stays instant.
const extractKey = (o) =>
[o.cavityErode, o.tongueReject, o.blobGrow, o.topBias, o.teethVerts].join('|');
/* ---------- build ---------- */
function makeXform(stab, neutral) {
@ -132,9 +166,10 @@ function makeXform(stab, neutral) {
function rebuild(resetKeep) {
if (!state.dense) return;
const o = opts();
state.lead = o.lead;
const N = state.dense.length;
state.stab = stabilize(state.dense, o.smoothWin);
state.stab = stabilize(state.dense, o.smoothWin, state.aspect);
const ap = state.stab.aperture;
const head = Math.max(1, Math.floor(N / 4));
@ -153,6 +188,12 @@ function rebuild(resetKeep) {
const apMax = Math.max(...ap);
state.hidden = ap.map((v) => v / apMax < o.apertureThresh);
if (state.images.length && state.extractKey !== extractKey(o)) {
state.interior = measureAll(state.images, state.dense, o);
state.extractKey = extractKey(o);
}
state.teeth = resolveTeeth(o);
// Plate outline per frame, so a kept frame shows its own head shape.
state.plates = state.stab.oval.map((r) => r.map(state.xform));
@ -179,6 +220,53 @@ function faceBoxes() {
});
}
// Presence gets hysteresis and a minimum dwell, the same treatment plate
// selection gets: a teeth block that blinks on and off for single frames is
// worse than one that is simply absent. Appearing needs a clear signal, staying
// needs only a weak one.
function resolveTeeth(o) {
const N = state.dense.length;
if (!state.interior) return new Array(N).fill({ show: false, pts: null });
const raw = state.interior.map((m, f) =>
(state.hidden[f] || !m.contour ? 0 : m.contrast));
const on = o.teethOn, off = o.teethOn * 0.7;
const shown = new Array(N).fill(false);
let live = false, since = 0;
for (let f = 0; f < N; f++) {
const want = live ? raw[f] > off : raw[f] > on;
if (want !== live && since >= o.teethDwell) { live = want; since = 0; }
else since++;
shown[f] = live && !state.hidden[f] && !!state.interior[f].contour;
}
// Into raster space through the same chain the lips take, including the
// isotropic aspect conversion - a contour in MediaPipe's normalised space is
// in the same stretched coordinates the landmarks are.
const toRaster = (pts, f) => {
const tf = state.stab.transforms[f];
return pts.map((p) => state.xform(applySim(tf, { x: p.x * state.aspect, y: p.y })));
};
const rast = state.interior.map((m, f) => (m.contour ? toRaster(m.contour, f) : null));
// Radial sampling makes vertex k mean the same direction on every frame, so
// smoothing across time is well defined and cannot reorder anything.
const sm = rast.map((pts, f) => {
if (!pts || !shown[f]) return pts;
const acc = pts.map(() => ({ x: 0, y: 0 }));
let c = 0;
for (let j = f - o.teethSmooth; j <= f + o.teethSmooth; j++) {
const k = Math.min(N - 1, Math.max(0, j));
if (!shown[k] || !rast[k] || rast[k].length !== pts.length) continue;
for (let v = 0; v < pts.length; v++) { acc[v].x += rast[k][v].x; acc[v].y += rast[k][v].y; }
c++;
}
return c ? acc.map((p) => ({ x: p.x / c, y: p.y / c })) : pts;
});
return shown.map((show, f) => ({ show, pts: sm[f] }));
}
/* ---------- render ---------- */
// The plate layer has several representations because its job changes: a flat
@ -197,8 +285,13 @@ function renderFrame(f, mode = plateMode()) {
if (mode === 'oval' || mode === 'oval+photo') r.fillPoly(state.plates[pf], IDX.base);
}
r.fillPoly(state.outer[f], IDX.dark); // mouth keeps every frame
if (!state.hidden[f]) r.fillPoly(state.inner[f], IDX.mouth);
const mf = leadIndex(f); // performance frame, possibly ahead
r.fillPoly(state.outer[mf], IDX.dark); // mouth keeps every frame
if (!state.hidden[mf]) {
r.fillPoly(state.inner[mf], IDX.mouth);
const te = state.teeth[mf];
if (te.show && te.pts && te.pts.length >= 3) r.fillPoly(te.pts, IDX.teeth);
}
return r;
}
@ -243,6 +336,22 @@ function compositeRender(canvas, f, zoom) {
const keptSorted = () => [...state.keep].sort((a, b) => a - b);
// Performance tracks can lead the clock.
//
// A centred moving average has no phase lag, so smoothing does not literally
// delay anything - but it blurs onsets, and the visually salient moment of a
// mouth opening moves later even though the mean does not. Animators also draw
// mouth shapes one or two frames ahead of the sound as a matter of course, so
// this is the normal control rather than a correction.
//
// Positive lead = the mouth arrives earlier. Only performance parts shift; the
// head stays with the audio, because it is the mouth that should anticipate.
function leadIndex(f) {
// Reads a cached scalar, not opts(): this runs once per strip thumbnail, and
// calling opts() here meant ~14 DOM reads x 74 frames on every redraw.
return shiftIndex(f, state.lead, state.dense.length);
}
function blit(canvas, raster, zoom) {
canvas.width = RW * zoom; canvas.height = RH * zoom;
canvas.getContext('2d').putImageData(raster.toImageData(PALETTE.map((p) => p.hex), zoom), 0, 0);
@ -258,8 +367,12 @@ function drawAll() {
function drawReadout() {
const kept = keptSorted();
const runs = kept.map((k, i) => (i + 1 < kept.length ? kept[i + 1] : state.dense.length) - k);
const lead = state.lead;
const teethFrames = state.teeth ? state.teeth.filter((t) => t.show).length : 0;
el('readout').textContent =
`${state.dense.length} frames → ${kept.length} drawings · ` +
`teeth on ${teethFrames}f · ` +
(lead ? `mouth leads ${lead}f (${(lead / state.fps * 1000).toFixed(0)}ms) · ` : '') +
`holds ${Math.min(...runs)}–${Math.max(...runs)} frames · ` +
`neutral f${state.neutral} · residual ` +
`${(state.stab.residual.reduce((a, b) => a + b, 0) / state.dense.length).toFixed(4)}`;
@ -268,8 +381,13 @@ function drawReadout() {
function drawPanes() {
const f = state.frame, kept = keptSorted();
const pf = heldFrame(kept, f);
// The mouth frame is always shown, not only when shifted, so the number can be
// watched diverging from f rather than taken on trust.
const lead = state.lead;
el('framelabel').textContent =
`f ${f} / ${state.dense.length - 1} · ${(f / state.fps).toFixed(2)}s · plate f${pf}` +
`f ${f} / ${state.dense.length - 1} · ${(f / state.fps).toFixed(2)}s · ` +
`plate f${pf} · mouth f${leadIndex(f)}` +
(lead ? ` (${lead > 0 ? '+' : ''}${lead} = ${(lead / state.fps * 1000).toFixed(0)}ms)` : '') +
(state.keep.has(f) ? ' · KEPT' : ' · held');
const c1 = el('cv-source'), g1 = c1.getContext('2d');
@ -301,11 +419,43 @@ function drawPanes() {
g2.moveTo(c2.width / 2, 0); g2.lineTo(c2.width / 2, c2.height);
g2.moveTo(0, c2.height / 2); g2.lineTo(c2.width, c2.height / 2); g2.stroke();
const z = (pts) => pts.map((p) => ({ x: p.x * ZOOM, y: p.y * ZOOM }));
const mf = leadIndex(f);
strokePts(g2, z(state.plates[pf]), '#3b4a63');
strokePts(g2, z(state.outer[f]), '#4ade80');
if (!state.hidden[f]) strokePts(g2, z(state.inner[f]), '#f87171');
// With a lead set, the unshifted contour is drawn as a ghost so the offset is
// something you can see rather than something you have to trust.
if (mf !== f) strokePts(g2, z(state.outer[f]), '#2f6b46');
strokePts(g2, z(state.outer[mf]), '#4ade80');
if (!state.hidden[mf]) strokePts(g2, z(state.inner[mf]), '#f87171');
compositeRender(el('cv-render'), f, ZOOM);
drawInteriorDebug(f);
}
// What the teeth measurement actually saw: sampled region, pixels above
// threshold in green, the resolved line in amber. Recomputed for the current
// frame only, so it costs nothing to keep on screen.
function drawInteriorDebug(fRaw) {
const f = leadIndex(fRaw);
const host = el('cv-teeth');
const img = state.images[f];
if (!img || state.hidden[f]) {
host.innerHTML = '';
el('teethinfo').textContent = state.images.length ? 'mouth closed' : 'no source frames';
return;
}
const o = opts();
const ctx = document.createElement('canvas').getContext('2d', { willReadFrequently: true });
const m = extractTeeth(img, LIPS_INNER.map((k) => state.dense[f][k]), ctx, o, true);
host.innerHTML = '';
if (m.debug) {
m.debug.style.width = '170px';
m.debug.style.imageRendering = 'pixelated';
host.append(m.debug);
}
const te = state.teeth[f];
el('teethinfo').textContent =
`contrast ${m.contrast.toFixed(3)} / gate ${o.teethOn.toFixed(2)} · ` +
`area ${m.area}px · ${te.show ? 'SHOWN' : 'hidden'}`;
}
function strokePts(g, pts, color, lw = 1) {
@ -400,11 +550,21 @@ function exportTake() {
// Sparse: one plate key per frame a human draws.
{ name: 'head', kind: 'plate', z: 0, interp: 'hold',
keys: kept.map((f, i) => ({ f: i, plate: i, src: f })) },
// Dense: the traced mouth keeps every frame.
// Dense: the traced mouth keeps every frame. The lead is baked in here -
// key f carries the pose from source frame f+lead - so the renderer never
// needs to know about it.
{ name: 'mouth', kind: 'poly', z: 30, color: 'skin_dark', interp: 'hold',
keys: state.outer.map((pts, f) => ({ f, src: f, pts })) },
keys: state.outer.map((_, f) => ({ f, src: leadIndex(f), pts: state.outer[leadIndex(f)] })) },
{ name: 'mouth_in', kind: 'poly', z: 31, color: 'mouth_dark', interp: 'hold', parent: 'mouth',
keys: state.inner.map((pts, f) => (state.hidden[f] ? { f, hidden: true } : { f, src: f, pts })) },
keys: state.inner.map((_, f) => {
const m = leadIndex(f);
return state.hidden[m] ? { f, hidden: true } : { f, src: m, pts: state.inner[m] };
}) },
{ name: 'teeth', kind: 'poly', z: 32, color: 'teeth', interp: 'hold', parent: 'mouth_in',
keys: state.teeth.map((_, f) => {
const m = leadIndex(f), te = state.teeth[m];
return te.show && te.pts ? { f, src: m, pts: te.pts } : { f, hidden: true };
}) },
],
};
const text = writeTake(take)
@ -440,11 +600,15 @@ async function runFrames() {
}
const { dense, missing } = await detectAll(images);
state.images = images; state.dense = dense;
state.aspect = images[0].naturalWidth / images[0].naturalHeight;
status('measuring mouth interiors…');
state.interior = measureAll(images, dense, opts());
el('scrub').max = dense.length - 1;
state.frame = 0;
rebuild(true);
const dur = (dense.length / state.fps).toFixed(2);
status(`${images.length} frames · ${state.fps}fps · ${dur}s` +
status(`${images.length} frames · ${images[0].naturalWidth}x${images[0].naturalHeight} · ` +
`${state.fps}fps · ${dur}s` +
(state.audio ? ' · audio loaded' : ' · no audio') +
(missing.length ? ` · no face on ${missing.length} (held previous)` : ''),
missing.length ? 'warn' : 'ok');
@ -455,6 +619,10 @@ function runSynthetic() {
state.images = [];
attachAudio(null);
state.fps = 12;
state.aspect = 1; // synthetic landmarks are generated square
state.lead = 0;
state.interior = null; // no pixels, so no teeth
state.dense = synthDense(72);
el('scrub').max = 71;
state.frame = 0;
@ -462,10 +630,16 @@ function runSynthetic() {
status('synthetic — exercises everything below detection', 'ok');
}
for (const id of ['verts', 'smoothWin', 'contourSmooth', 'apertureThresh', 'tol']) {
for (const id of ['verts', 'smoothWin', 'contourSmooth', 'apertureThresh', 'tol',
'teethOn', 'teethDwell', 'teethErode', 'tongueReject', 'blobGrow',
'topBias', 'teethVerts', 'teethSmooth', 'lead']) {
el(id).addEventListener('input', () => {
el(id + 'v').textContent = id === 'apertureThresh' || id === 'tol'
? (+el(id).value / 1000).toFixed(3) : el(id).value;
? (+el(id).value / 1000).toFixed(3)
: ['teethOn', 'teethErode', 'tongueReject', 'topBias'].includes(id)
? (+el(id).value / 100).toFixed(2)
: id === 'lead' && +el(id).value > 0 ? `+${el(id).value}`
: el(id).value;
if (id === 'tol') return; // tol only matters when you ask for a suggestion
rebuild(false);
});
@ -533,6 +707,12 @@ window.addEventListener('keydown', (e) => {
else if (e.key === 'ArrowLeft') { seekTo(Math.max(0, state.frame - 1)); drawAll(); }
else if (e.key === 'Backspace' || e.key === 'Delete' || e.key === 'x') {
state.keep.delete(state.frame === 0 ? -1 : state.frame); drawAll();
} else if (e.key === '[' || e.key === ']') {
const n = el('lead');
n.value = Math.max(+n.min, Math.min(+n.max, +n.value + (e.key === ']' ? 1 : -1)));
state.lead = +n.value;
el('leadv').textContent = state.lead > 0 ? `+${state.lead}` : String(state.lead);
drawAll();
} else if (e.key === 'b') {
const sel = el('plateMode');
sel.selectedIndex = (sel.selectedIndex + 1) % sel.options.length;
@ -585,6 +765,11 @@ PALETTE.forEach((p) => {
// #synth / #frames autorun, so the tool can be driven headlessly for smoke tests
// and deep-linked. Detection needs WebGL; the synthetic path does not.
window.addEventListener('error', (e) => {
const s = document.getElementById('status');
if (s) { s.textContent = e.message; s.className = 'err'; }
});
if (location.hash === '#synth') runSynthetic();
else if (location.hash === '#frames') runFrames();
else status('ready — Load frames, then step with \u2190 \u2192 and delete with X');

225
js/interior.js Normal file
View file

@ -0,0 +1,225 @@
// Mouth interior from image content.
//
// MediaPipe has no landmarks inside the lips - the inner ring bounds the cavity
// and everything within it is just pixels. So teeth come from the picture.
//
// The hazard is vertex correspondence. A traced contour reorders between frames
// and boils, which is the failure docs/design.md exists to avoid. The way
// out for a blob specifically is RADIAL SAMPLING: march outward from the
// centroid along N fixed directions and take the last pixel inside. Vertex k is
// then always "the blob's extent in direction k" - correspondence holds by
// construction, the vertex count is fixed, and the result smooths over time
// without any reordering being possible. It also yields a star-shaped
// reduction, which is what flat blocks of colour want anyway.
export function otsuForTest(h, t) { return otsu(h, t); }
// Otsu's threshold plus its two class means. The means matter as much as the
// threshold: Otsu ALWAYS returns a split, including on a homogeneous region, so
// their separation is the only thing that says the split means anything.
function otsu(hist, total) {
let sum = 0;
for (let i = 0; i < 256; i++) sum += i * hist[i];
let sumB = 0, wB = 0, best = 0, bestVar = -1, bestDark = 0, bestBright = 0;
for (let t = 0; t < 256; t++) {
wB += hist[t];
if (!wB) continue;
const wF = total - wB;
if (!wF) break;
sumB += t * hist[t];
const mDark = sumB / wB, mBright = (sum - sumB) / wF;
const between = wB * wF * (mDark - mBright) * (mDark - mBright);
if (between > bestVar) { bestVar = between; best = t; bestDark = mDark; bestBright = mBright; }
}
return { thr: best, mDark: bestDark, mBright: bestBright };
}
const pointInPoly = (pts, x, y) => {
let inside = false;
for (let i = 0, j = pts.length - 1; i < pts.length; j = i++) {
if ((pts[i].y > y) !== (pts[j].y > y) &&
x < ((pts[j].x - pts[i].x) * (y - pts[i].y)) / (pts[j].y - pts[i].y) + pts[i].x) inside = !inside;
}
return inside;
};
// Shrink or grow a ring about its centroid. MediaPipe's inner lip landmarks sit
// slightly OUTSIDE the real opening, so sampling the ring as given includes lip
// pixels - bright, and right at the boundary where they do most damage.
export function scaleRing(pts, k) {
let cx = 0, cy = 0;
for (const p of pts) { cx += p.x; cy += p.y; }
cx /= pts.length; cy /= pts.length;
return pts.map((p) => ({ x: cx + (p.x - cx) * k, y: cy + (p.y - cy) * k }));
}
/* ---- binary morphology on the candidate mask ---- */
function erodeMask(m, w, h) {
const o = new Uint8Array(m.length);
for (let y = 1; y < h - 1; y++) for (let x = 1; x < w - 1; x++) {
const i = y * w + x;
o[i] = m[i] && m[i - 1] && m[i + 1] && m[i - w] && m[i + w] ? 1 : 0;
}
return o;
}
function dilateMask(m, w, h) {
const o = new Uint8Array(m.length);
for (let y = 1; y < h - 1; y++) for (let x = 1; x < w - 1; x++) {
const i = y * w + x;
o[i] = m[i] || m[i - 1] || m[i + 1] || m[i - w] || m[i + w] ? 1 : 0;
}
return o;
}
// Largest 4-connected component, scored with a bias toward the TOP of the
// cavity: upper teeth hang from the lip, and the usual false positive is the
// tongue sitting lower down. Area alone picks the tongue when the mouth is wide.
function bestComponent(mask, w, h, topBias) {
const label = new Int32Array(mask.length).fill(-1);
const stack = [];
let best = null, id = 0;
for (let s = 0; s < mask.length; s++) {
if (!mask[s] || label[s] >= 0) continue;
stack.length = 0; stack.push(s);
label[s] = id;
const px = [];
let sumY = 0;
while (stack.length) {
const i = stack.pop();
px.push(i);
sumY += (i / w) | 0;
const x = i % w, y = (i / w) | 0;
if (x > 0 && mask[i - 1] && label[i - 1] < 0) { label[i - 1] = id; stack.push(i - 1); }
if (x < w - 1 && mask[i + 1] && label[i + 1] < 0) { label[i + 1] = id; stack.push(i + 1); }
if (y > 0 && mask[i - w] && label[i - w] < 0) { label[i - w] = id; stack.push(i - w); }
if (y < h - 1 && mask[i + w] && label[i + w] < 0) { label[i + w] = id; stack.push(i + w); }
}
const meanY = sumY / px.length / h; // 0 top, 1 bottom
const score = px.length * (1 - topBias * meanY);
if (!best || score > best.score) best = { score, px, area: px.length, meanY };
id++;
}
return best;
}
// Radial sampling from the centroid: N fixed directions, last pixel inside.
function radialContour(mask, w, h, cx, cy, n) {
const pts = [];
const maxR = Math.hypot(w, h);
let prev = 1;
for (let k = 0; k < n; k++) {
const a = -(k / n) * Math.PI * 2; // slot 0 = +x, 5/20 = top
const dx = Math.cos(a), dy = Math.sin(a);
let hit = 0;
for (let r = 0.5; r < maxR; r += 0.5) {
const x = Math.round(cx + dx * r), y = Math.round(cy + dy * r);
if (x < 0 || y < 0 || x >= w || y >= h) break;
if (mask[y * w + x]) hit = r;
else if (hit > 0 && r > hit + 2) break; // tolerate a 2px gap, then stop
}
// A ray that escapes immediately would collapse the polygon; hold the last
// good radius so the shape stays closed rather than spiking to the centre.
if (hit <= 0) hit = prev * 0.6;
prev = hit;
pts.push({ x: cx + dx * hit, y: cy + dy * hit });
}
return pts;
}
/* ---- the extraction ---- */
export function extractTeeth(img, innerNorm, ctx, o, wantDebug = false) {
const none = { contour: null, contrast: 0, area: 0, debug: null };
const ring = scaleRing(innerNorm, 1 - (o.cavityErode ?? 0.18));
let x0 = 1, y0 = 1, x1 = 0, y1 = 0;
for (const p of ring) {
x0 = Math.min(x0, p.x); y0 = Math.min(y0, p.y);
x1 = Math.max(x1, p.x); y1 = Math.max(y1, p.y);
}
const W = img.naturalWidth, H = img.naturalHeight;
const px0 = Math.max(0, Math.floor(x0 * W)), py0 = Math.max(0, Math.floor(y0 * H));
const pw = Math.min(W - px0, Math.ceil((x1 - x0) * W)), ph = Math.min(H - py0, Math.ceil((y1 - y0) * H));
if (pw < 5 || ph < 5) return none;
ctx.canvas.width = pw; ctx.canvas.height = ph;
ctx.drawImage(img, px0, py0, pw, ph, 0, 0, pw, ph);
const src = ctx.getImageData(0, 0, pw, ph);
const d = src.data;
const poly = ring.map((p) => ({ x: p.x * W - px0, y: p.y * H - py0 }));
const hist = new Uint32Array(256);
const lum = new Float32Array(pw * ph);
const red = new Float32Array(pw * ph);
const inReg = new Uint8Array(pw * ph);
let n = 0;
for (let y = 0; y < ph; y++) for (let x = 0; x < pw; x++) {
if (!pointInPoly(poly, x + 0.5, y + 0.5)) continue;
const i = y * pw + x, oo = i * 4;
const R = d[oo], G = d[oo + 1], B = d[oo + 2];
lum[i] = (0.299 * R + 0.587 * G + 0.114 * B) | 0;
// Tongue is red relative to its own brightness; teeth are near-neutral.
red[i] = (R - (G + B) / 2) / 255;
inReg[i] = 1; hist[lum[i]]++; n++;
}
if (n < 24) return none;
const { thr, mDark, mBright } = otsu(hist, n);
const contrast = (mBright - mDark) / 255;
let mask = new Uint8Array(pw * ph);
for (let i = 0; i < mask.length; i++) {
mask[i] = inReg[i] && lum[i] > thr && red[i] < (o.tongueReject ?? 0.18) ? 1 : 0;
}
// Open once to despeckle, then apply the signed size adjustment.
mask = dilateMask(erodeMask(mask, pw, ph), pw, ph);
const grow = o.blobGrow | 0;
for (let k = 0; k < Math.abs(grow); k++) {
mask = grow < 0 ? erodeMask(mask, pw, ph) : dilateMask(mask, pw, ph);
}
const comp = bestComponent(mask, pw, ph, o.topBias ?? 0.6);
if (!comp || comp.area < (o.minArea ?? 12)) {
return { contour: null, contrast, area: comp ? comp.area : 0,
debug: wantDebug ? debugCanvas(src, inReg, mask, pw, ph, null) : null };
}
const only = new Uint8Array(mask.length);
let cx = 0, cy = 0;
for (const i of comp.px) { only[i] = 1; cx += i % pw; cy += (i / pw) | 0; }
cx /= comp.px.length; cy /= comp.px.length;
const local = radialContour(only, pw, ph, cx, cy, o.teethVerts ?? 10);
const contour = local.map((p) => ({ x: (p.x + px0) / W, y: (p.y + py0) / H }));
return {
contour, contrast, area: comp.area,
debug: wantDebug ? debugCanvas(src, inReg, only, pw, ph, local) : null,
};
}
// Sampled region dimmed, kept pixels green, extracted contour in amber.
function debugCanvas(src, inReg, mask, pw, ph, local) {
const c = document.createElement('canvas');
c.width = pw; c.height = ph;
const g = c.getContext('2d');
const out = new ImageData(pw, ph);
for (let i = 0; i < pw * ph; i++) {
const o = i * 4;
const [r, gr, b] = [src.data[o], src.data[o + 1], src.data[o + 2]];
if (!inReg[i]) { out.data[o] = r * 0.25; out.data[o + 1] = gr * 0.25; out.data[o + 2] = b * 0.25; }
else if (mask[i]) { out.data[o] = 60; out.data[o + 1] = 230; out.data[o + 2] = 120; }
else { out.data[o] = r; out.data[o + 1] = gr; out.data[o + 2] = b; }
out.data[o + 3] = 255;
}
g.putImageData(out, 0, 0);
if (local && local.length) {
g.strokeStyle = '#fbbf24'; g.lineWidth = 1;
g.beginPath();
local.forEach((p, i) => (i ? g.lineTo(p.x, p.y) : g.moveTo(p.x, p.y)));
g.closePath(); g.stroke();
}
return c;
}

View file

@ -1,7 +1,7 @@
// MediaPipe FaceLandmarker index tables.
// Ring arrays are ORDERED traversals, not raw connection sets: vertex position
// within a ring is the vertex's identity, and every downstream stage depends on
// that ordering being stable. See docs/roto-puppet.md, "Fixed topology".
// that ordering being stable. See docs/design.md, "Fixed topology".
// Rigid landmarks for the similarity fit. Eye corners, nose bridge, nose tip.
// Nothing here may be a feature that moves under performance: including the

View file

@ -1,17 +1,26 @@
// Analysis: dense track -> stabilised head-local contours -> selected keys.
// All policy lives here, never in the renderer. See docs/roto-puppet.md,
// All policy lives here, never in the renderer. See docs/design.md,
// "The take is the contract".
import { RIGID, LIPS_OUTER, LIPS_INNER, APERTURE, FACE_OVAL, EYE_INNER, subsampleSlots } from './landmarks.js';
import { fitSimilarity, applySimAll, applySim, fitResidual, procrustesMean, smoothTransforms, movingAverage } from './mathutil.js';
const pick = (lm, idx) => idx.map((i) => ({ x: lm[i].x, y: lm[i].y }));
// MediaPipe normalises x by image WIDTH and y by image HEIGHT, so its normalised
// space is anisotropic: for a 1080x1920 frame, one unit of x is 1080px and one
// unit of y is 1920px. Treating those as comparable stretches everything
// horizontally by H/W, and worse, makes fitSimilarity fit a "rotation" in a
// sheared space, so head roll comes out subtly wrong as well.
//
// Multiplying x by aspect = W/H converts to an ISOTROPIC space whose unit is one
// image height, so equal numbers mean equal pixels. Everything downstream -
// Procrustes, the similarity fit, the raster transform - depends on that.
const pick = (lm, idx, aspect) => idx.map((i) => ({ x: lm[i].x * aspect, y: lm[i].y }));
// Stage 1-3: fit the rigid transform per frame, smooth its parameters, then map
// every contour through it into the reference frame. The result is head-local:
// translation, roll and depth-scale of the head are gone.
export function stabilize(dense, smoothRadius) {
const rigid = dense.map((f) => pick(f, RIGID));
export function stabilize(dense, smoothRadius, aspect = 1) {
const rigid = dense.map((f) => pick(f, RIGID, aspect));
const ref = procrustesMean(rigid);
const raw = rigid.map((r) => fitSimilarity(r, ref));
const tfs = smoothTransforms(raw, smoothRadius);
@ -26,12 +35,12 @@ export function stabilize(dense, smoothRadius) {
// Residual rises with out-of-plane rotation, which no 2D similarity can
// remove. High values mean this section wants a different head plate.
residual: tfs.map((tf, i) => fitResidual(tf, rigid[i], ref)),
outer: dense.map((f, i) => applySimAll(tfs[i], pick(f, LIPS_OUTER))),
inner: dense.map((f, i) => applySimAll(tfs[i], pick(f, LIPS_INNER))),
oval: dense.map((f, i) => applySimAll(tfs[i], pick(f, FACE_OVAL))),
eyes: dense.map((f, i) => applySimAll(tfs[i], pick(f, EYE_INNER))),
outer: dense.map((f, i) => applySimAll(tfs[i], pick(f, LIPS_OUTER, aspect))),
inner: dense.map((f, i) => applySimAll(tfs[i], pick(f, LIPS_INNER, aspect))),
oval: dense.map((f, i) => applySimAll(tfs[i], pick(f, FACE_OVAL, aspect))),
eyes: dense.map((f, i) => applySimAll(tfs[i], pick(f, EYE_INNER, aspect))),
aperture: dense.map((f, i) => {
const a = applySimAll(tfs[i], pick(f, APERTURE));
const a = applySimAll(tfs[i], pick(f, APERTURE, aspect));
return Math.hypot(a[0].x - a[1].x, a[0].y - a[1].y);
}),
};
@ -113,7 +122,7 @@ export function activeKey(keys, f) {
// Temporal smoothing of a contour, per vertex, across time.
//
// docs/roto-puppet.md says to smooth the transform and never the contour. That
// docs/design.md says to smooth the transform and never the contour. That
// was correct while keys were sparse: sampling at velocity minima rejected
// per-frame detector noise for free. With a key on every frame the noise is
// visible as a shimmer along the lip edge, so a bounded exception applies -
@ -168,3 +177,12 @@ export function heldFrame(kept, f) {
for (const k of kept) { if (k <= f) hit = k; else break; }
return hit;
}
// Shift a performance track against the clock, clamped at the ends.
//
// Pure and exported so the shift can actually be asserted: "the slider feels
// like it does nothing" is otherwise indistinguishable from "the slider does
// nothing", and at 24fps a lead of 1 is 42ms, which is small enough to doubt.
export function shiftIndex(f, lead, n) {
return Math.min(n - 1, Math.max(0, f + lead));
}

View file

@ -2,16 +2,17 @@
// module graph the tool uses is what gets tested.
//
// The ring-simplicity check exists because "fixed topology" is load-bearing in
// docs/roto-puppet.md: because hold parts CUT between poses rather than
// docs/design.md: because hold parts CUT between poses rather than
// interpolating, a ring whose vertex order is wrong self-intersects and renders
// as blocks meeting at corners. It is invisible at some vertex counts and obvious
// at others, so it needs an assertion rather than an eyeball.
import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, subsampleSlots, subsampleRing } from './landmarks.js';
import { fitSimilarity, applySim, procrustesMean, smoothTransforms } from './mathutil.js';
import { stabilize, toRasterRing, selectKeys, activeKey } from './pipeline.js';
import { stabilize, toRasterRing, selectKeys, activeKey, shiftIndex } from './pipeline.js';
import { IndexedRaster, hexToRgb } from './raster.js';
import { writeTake } from './take.js';
import { otsuForTest, scaleRing } from './interior.js';
import { synthDense } from './synth.js';
const results = [];
@ -44,6 +45,50 @@ const spreadX = (frames, slot) => {
/* ---- the tests ---- */
// Cross-check every el('id') in app.js against the ids in index.html.
//
// This bug class has bitten twice: a knob wired in app.js but absent from the
// markup throws during wiring, which aborts the rest of the module and leaves a
// blank page. The symptom ("nothing happens") points nowhere near the cause, so
// it is worth an automated check rather than vigilance.
export async function runWiring() {
const out = [];
try {
const [app, html] = await Promise.all([
fetch('./js/app.js').then((r) => r.text()),
fetch('./index.html').then((r) => r.text()),
]);
const ids = new Set([...app.matchAll(/\bel\(\s*['"]([\w-]+)['"]\s*\)/g)].map((m) => m[1]));
const list = app.match(/for \(const id of \[([\s\S]*?)\]\)/);
if (list) {
for (const m of list[1].matchAll(/'([\w]+)'/g)) { ids.add(m[1]); ids.add(m[1] + 'v'); }
}
// Duplicate keys in an object literal are silent in JS - the last one wins.
// In opts() that meant the teeth vertex slider was quietly driving the lip
// vertex count while the lip slider did nothing at all.
const lit = app.match(/const opts = \(\) => \(\{([\s\S]*?)\n\}\);/);
if (lit) {
const keys = [...lit[1].matchAll(/^\s*([A-Za-z_$][\w$]*)\s*:/gm)].map((m) => m[1]);
const dupes = keys.filter((k, i) => keys.indexOf(k) !== i);
out.push({ name: `opts() has no duplicate keys (${keys.length} checked)`,
pass: dupes.length === 0, detail: [...new Set(dupes)].join(', ') });
} else {
out.push({ name: 'opts() literal found for duplicate-key check', pass: false, detail: '' });
}
const have = new Set([...html.matchAll(/id="([\w-]+)"/g)].map((m) => m[1]));
const missing = [...ids].filter((i) => !have.has(i));
out.push({ name: `every el() id exists in index.html (${ids.size} checked)`,
pass: missing.length === 0, detail: missing.join(', ') });
const unused = [...have].filter((i) => !ids.has(i));
out.push({ name: 'no orphaned ids in index.html', pass: unused.length === 0,
detail: unused.join(', ') });
} catch (e) {
out.push({ name: 'wiring check ran', pass: false, detail: e.message });
}
return out;
}
export function run() {
results.length = 0;
@ -113,6 +158,40 @@ export function run() {
const apRange = Math.max(...stab.aperture) - Math.min(...stab.aperture);
ok('stabilisation preserves mouth motion', apRange > 0.05, `aperture range ${apRange.toFixed(4)}`);
// ASPECT: a shape that is circular in PIXEL space must stay circular in raster
// space. MediaPipe normalises x by width and y by height, so for a portrait
// frame equal normalised numbers are unequal pixel distances; feeding those
// straight through stretches everything horizontally by H/W. This asserts the
// isotropic conversion, and fails at ~1.78 for a 1080x1920 clip without it.
for (const [W, H] of [[1080, 1920], [1920, 1080], [640, 640]]) {
const aspect = W / H;
const N = 24, cx = 0.5, cy = 0.5, rPx = 200;
// a true circle of radius rPx, expressed in MediaPipe normalised coords
const circleFrames = [];
for (let t = 0; t < 4; t++) {
const pts = new Array(478).fill(null).map(() => ({ x: 0.5, y: 0.5, z: 0 }));
RIGID.forEach((id, k) => {
const a = (k / RIGID.length) * Math.PI * 2;
pts[id] = { x: cx + (120 * Math.cos(a)) / W, y: cy + (120 * Math.sin(a)) / H, z: 0 };
});
LIPS_OUTER.forEach((id, k) => {
const a = -(k / LIPS_OUTER.length) * Math.PI * 2;
pts[id] = { x: cx + (rPx * Math.cos(a)) / W, y: cy + (rPx * Math.sin(a)) / H, z: 0 };
});
FACE_OVAL.forEach((id, k) => {
const a = -(k / FACE_OVAL.length) * Math.PI * 2;
pts[id] = { x: cx + (420 * Math.cos(a)) / W, y: cy + (420 * Math.sin(a)) / H, z: 0 };
});
circleFrames.push(pts);
}
const st2 = stabilize(circleFrames, 0, aspect);
const ring = toRasterRing(st2.outer[0], LIPS_OUTER, 16, (p) => p);
const xs = ring.map((p) => p.x), ys = ring.map((p) => p.y);
const ratio = (Math.max(...xs) - Math.min(...xs)) / (Math.max(...ys) - Math.min(...ys));
ok(`circle stays circular at ${W}x${H}`, Math.abs(ratio - 1) < 0.02,
`w/h ratio ${ratio.toFixed(4)}`);
}
// key selection
const xf = (p) => ({ x: p.x * 320, y: p.y * 200 });
const shapes = stab.outer.map((r) => toRasterRing(r, LIPS_OUTER, 8, xf));
@ -127,6 +206,31 @@ export function run() {
ok('activeKey holds between keys',
activeKey(sel.keys, sel.keys[1].f - 1).f === sel.keys[0].f);
// mouth lead: a shift that "feels like it does nothing" is indistinguishable
// from one that does nothing, so assert the arithmetic directly.
ok('lead 0 is identity', [0, 5, 71].every((f) => shiftIndex(f, 0, 72) === f));
ok('positive lead moves the source frame forward', shiftIndex(10, 2, 72) === 12);
ok('negative lead moves it back', shiftIndex(10, -3, 72) === 7);
ok('lead clamps at the start', shiftIndex(1, -6, 72) === 0);
ok('lead clamps at the end', shiftIndex(70, 6, 72) === 71);
{
// ...and that it selects different POSES, not merely different indices.
// Checked across the whole track rather than at one pair: synthetic poses
// hold for nine-frame beats, so any single pair can legitimately be
// identical while the shift works perfectly.
const N = shapes.length;
let moved = 0, total = 0;
for (let f = 0; f < N; f++) {
const a = shapes[shiftIndex(f, 0, N)], b = shapes[shiftIndex(f, 3, N)];
let d = 0;
for (let i = 0; i < a.length; i++) d += Math.hypot(a[i].x - b[i].x, a[i].y - b[i].y);
total++;
if (d / a.length > 0.5) moved++;
}
ok('a lead of 3 changes the pose on a good share of frames', moved / total > 0.2,
`${moved}/${total} frames differ`);
}
// rasteriser: indexed, hard-edged, no blending
const r = new IndexedRaster(64, 48);
r.clear(0);
@ -147,6 +251,43 @@ export function run() {
ok('palette expansion introduces no intermediate colours',
[...seen].every((c) => allowed.has(c)), `${seen.size} distinct colours`);
// scaleRing is what pulls the sampled region in from MediaPipe's inner lip
// landmarks, which sit slightly outside the real opening.
{
const ring = [{ x: 0, y: 0 }, { x: 10, y: 0 }, { x: 10, y: 10 }, { x: 0, y: 10 }];
const small = scaleRing(ring, 0.5);
const w = Math.max(...small.map((p) => p.x)) - Math.min(...small.map((p) => p.x));
ok('scaleRing(0.5) halves the extent', Math.abs(w - 5) < 1e-9, `width ${w}`);
const same = scaleRing(ring, 1);
ok('scaleRing(1) is identity', same.every((p, i) => Math.abs(p.x - ring[i].x) < 1e-9));
let cx = 0; for (const p of small) cx += p.x;
ok('scaleRing keeps the centroid', Math.abs(cx / 4 - 5) < 1e-9);
}
// Otsu on a uniform region must report near-zero class separation. It will
// still return a threshold - that is what Otsu does - so the separation is the
// only thing that distinguishes "found teeth" from "split noise in a dark
// mouth", which is what made the band fill the whole cavity.
{
const flat = new Uint32Array(256); flat[40] = 500;
const f = otsuForTest(flat, 500);
ok('uniform region yields ~no class separation',
Math.abs(f.mBright - f.mDark) / 255 < 0.02, `sep ${((f.mBright - f.mDark) / 255).toFixed(4)}`);
const noisy = new Uint32Array(256);
for (let i = 30; i <= 60; i++) noisy[i] = 20; // dark cavity, some spread
const nz = otsuForTest(noisy, 31 * 20);
ok('dark-but-noisy region stays below a sane gate',
(nz.mBright - nz.mDark) / 255 < 0.14, `sep ${((nz.mBright - nz.mDark) / 255).toFixed(4)}`);
const teeth = new Uint32Array(256);
for (let i = 20; i <= 45; i++) teeth[i] = 40; // cavity
for (let i = 180; i <= 220; i++) teeth[i] = 30; // teeth
const tt = otsuForTest(teeth, 26 * 40 + 41 * 30);
ok('real bright/dark split clears the gate',
(tt.mBright - tt.mDark) / 255 > 0.4, `sep ${((tt.mBright - tt.mDark) / 255).toFixed(4)}`);
}
// take writer round-trip
const take = {
name: 'test', frames: 72, width: 320, height: 200, exposure: 2,

View file

@ -1,4 +1,4 @@
// Take-file writer. Format is specified in docs/roto-puppet.md, "The take
// Take-file writer. Format is specified in docs/design.md, "The take
// format". Line-oriented text on purpose: Poco can parse it with fopen/fgets
// from poco/src/safefile.c and strtok/atoi/atof from poco/src/strlib.c, so the
// Animator Pro render script needs no new native code.

View file

@ -10,11 +10,13 @@ import { applySim } from './mathutil.js';
// Compose pixel-space -> raster-space into one affine.
//
// MediaPipe normalises x by width and y by height, so normalised space is a
// stretched pixel space and the composition is a general affine rather than a
// similarity. Three mapped points determine it exactly.
// Landmarks are converted to an isotropic space (unit = one image height) before
// fitting, so pixels map in the same way: BOTH axes divide by imgH, not by their
// own dimension. Dividing x by imgW here instead is what stretched the underlay
// horizontally by H/W and made it disagree with nothing - it matched the equally
// wrong vector shapes.
export function frameAffine(tf, xform, imgW, imgH) {
const map = (px, py) => xform(applySim(tf, { x: px / imgW, y: py / imgH }));
const map = (px, py) => xform(applySim(tf, { x: px / imgH, y: py / imgH }));
const P0 = map(0, 0), P1 = map(imgW, 0), P2 = map(0, imgH);
return {
a: (P1.x - P0.x) / imgW, b: (P1.y - P0.y) / imgW,

View file

@ -1 +1 @@
{"fps":12,"frames":37,"dir":"frames","audio":"audio.wav","source":"IMG_8486.MOV"}
{"fps":24,"frames":74,"dir":"frames","audio":"audio.wav","source":"IMG_8486.MOV"}

View file

@ -1,79 +0,0 @@
<!doctype html><html><head><meta charset="utf-8"><title>probe…</title></head>
<body><pre id="o">running…</pre>
<script type="module">
import { FaceLandmarker, FilesetResolver } from 'https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@1.0.1/vision_bundle.mjs';
import { LIPS_OUTER, LIPS_INNER, subsampleSlots } from './js/landmarks.js';
import { stabilize, toRasterRing, selectKeys } from './js/pipeline.js';
const log = [];
const say = (s) => { log.push(s); document.getElementById('o').textContent = log.join('\n'); };
function crosses(a,b,c,d){const o=(p,q,r)=>Math.sign((q.x-p.x)*(r.y-p.y)-(q.y-p.y)*(r.x-p.x));
const o1=o(a,b,c),o2=o(a,b,d),o3=o(c,d,a),o4=o(c,d,b);
return o1!==o2&&o3!==o4&&o1!==0&&o2!==0&&o3!==0&&o4!==0;}
function selfInts(pts){const n=pts.length,h=[];for(let i=0;i<n;i++)for(let j=i+1;j<n;j++){
if((j+1)%n===i||(i+1)%n===j)continue;
if(crosses(pts[i],pts[(i+1)%n],pts[j],pts[(j+1)%n]))h.push([i,j]);}return h;}
const loadImg = (src) => new Promise(r => { const i=new Image(); i.onload=()=>r(i); i.onerror=()=>r(null); i.src=src; });
try {
const imgs=[];
for(let i=1;i<=900;i++){const im=await loadImg(`frames/${String(i).padStart(4,'0')}.png`); if(!im)break; imgs.push(im);}
say(`frames loaded: ${imgs.length} @ ${imgs[0].naturalWidth}x${imgs[0].naturalHeight}`);
const fs = await FilesetResolver.forVisionTasks('https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@1.0.1/wasm');
const lm = await FaceLandmarker.createFromOptions(fs, {
baseOptions:{ modelAssetPath:'./face_landmarker.task', delegate:'CPU' },
runningMode:'IMAGE', numFaces:1 });
say('landmarker ready (CPU delegate)');
const cv=document.createElement('canvas'); const dense=[]; let miss=0;
for(const im of imgs){
cv.width=im.naturalWidth; cv.height=im.naturalHeight;
cv.getContext('2d').drawImage(im,0,0);
const out=lm.detect(cv);
if(out.faceLandmarks?.length) dense.push(out.faceLandmarks[0]);
else { miss++; if(dense.length) dense.push(dense[dense.length-1]); }
}
say(`detected: ${dense.length}/${imgs.length} (no face on ${miss})`);
if(!dense.length) throw new Error('no face detected in any frame');
// THE key check: are LIPS_OUTER / LIPS_INNER correct traversals of real data?
for(const [name,tab] of [['LIPS_OUTER',LIPS_OUTER],['LIPS_INNER',LIPS_INNER]]){
let bad=0, first=null;
for(let n=4;n<=16;n+=2){
const slots=subsampleSlots(tab.length,n);
for(let f=0;f<dense.length;f++){
const h=selfInts(slots.map(s=>dense[f][tab[s]]));
if(h.length){bad++; first=first||`verts=${n} f=${f} edges ${JSON.stringify(h[0])}`;}
}
}
say(`${name}: ${bad===0?'SIMPLE at every budget/frame':`SELF-INTERSECTS ${bad}x first ${first}`}`);
// full 20-ring too
let bad20=0;
for(let f=0;f<dense.length;f++) if(selfInts(tab.map(i=>dense[f][i])).length) bad20++;
say(` full 20-point ring: ${bad20===0?'simple on all frames':`self-intersects on ${bad20} frames`}`);
}
const st = stabilize(dense, 5);
const res = st.residual;
const mean = res.reduce((a,b)=>a+b,0)/res.length;
say(`residual mean ${mean.toFixed(5)} max ${Math.max(...res).toFixed(5)} (high = out-of-plane rotation)`);
const eyeX = st.eyes.map(e=>e[0].x);
const rawX = dense.map(f=>f[133].x);
say(`eye-inner x spread: raw ${(Math.max(...rawX)-Math.min(...rawX)).toFixed(4)} -> stabilised ${(Math.max(...eyeX)-Math.min(...eyeX)).toFixed(4)}`);
const ap=st.aperture;
say(`aperture min ${Math.min(...ap).toFixed(4)} max ${Math.max(...ap).toFixed(4)} range ${(Math.max(...ap)-Math.min(...ap)).toFixed(4)}`);
let nf=0; ap.forEach((v,i)=>{ if(v===Math.min(...ap)) nf=i; });
say(`most-closed frame: ${nf}`);
const xf=p=>({x:p.x*320,y:p.y*200});
const shapes=st.outer.map(r=>toRasterRing(r,LIPS_OUTER,8,xf));
for(const [mh,dt] of [[1,0.6],[2,0.6],[2,1.5],[2,3.0],[3,1.5]]){
const k=selectKeys(shapes,{minHold:mh,distThresh:dt,velSmooth:3,exposure:1});
say(`minHold=${mh} gate=${dt}: ${k.candidates.length} cand -> ${k.keys.length} keys [${k.keys.map(x=>x.f).join(' ')}]`);
}
document.title='PROBE OK';
} catch(e){ say('ERROR: '+e.message+'\n'+e.stack); document.title='PROBE FAIL'; }
</script></body></html>

View file

@ -7,10 +7,11 @@
</style></head><body>
<h1 id="head">running…</h1><ul id="out"></ul>
<script type="module">
import { run } from './js/selftest.js';
import { run, runWiring } from './js/selftest.js';
let res;
try { res = run(); }
catch (e) { res = [{ name: 'harness threw: ' + e.message, pass: false, detail: String(e.stack).split('\n')[1] || '' }]; }
res = res.concat(await runWiring());
const pass = res.filter(r => r.pass).length;
const fail = res.length - pass;
document.getElementById('head').textContent = `${fail === 0 ? 'PASS' : 'FAIL'} — ${pass}/${res.length} assertions`;