Compare commits
10 commits
de90bfd907
...
c2603102da
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
c2603102da | ||
| f827757f7d | |||
| dc0a3f34e6 | |||
| f9b8ec8617 | |||
| d18791e6ad | |||
| a146d271b6 | |||
| 9a11eeabc1 | |||
| c6daf827d1 | |||
| 941022b69f | |||
| 29b0bd6690 |
14 changed files with 921 additions and 137 deletions
73
README.md
73
README.md
|
|
@ -1,16 +1,18 @@
|
||||||
# roto
|
# arthur
|
||||||
|
|
||||||
Video → **take file** builder for the Animator Pro rotoscope/puppet pipeline
|
An animation suite for turning live-action video into work that reads as
|
||||||
described in `../docs/roto-puppet.md`.
|
hand-authored — flat shapes, a tiny palette, hard edges, motion carried by
|
||||||
|
silhouette, in the idiom of *Another World*.
|
||||||
|
|
||||||
This is the analysis and tuning half. It stabilises a face out of a clip, reduces
|
It stabilises a face out of a clip, reduces the lip contour to a handful of
|
||||||
the lip contour to a handful of vertices, selects sparse keys on motion extremes,
|
vertices, derives teeth from image content, lets you decide which frames need
|
||||||
and previews the result as flat indexed fills — so the *look and timing* can be
|
their own hand-drawn head, and renders the result as flat indexed fills at
|
||||||
judged in seconds rather than through a minutes-long Animator Pro render. It
|
320×200 with no antialiasing.
|
||||||
emits a `.take` file; nothing here touches Animator Pro.
|
|
||||||
|
|
||||||
All policy lives here. The take arrives at the renderer with its keys already
|
The constraint is deliberate. 320×200 and an indexed palette are inherited from
|
||||||
chosen.
|
Animator Pro, where this started, but they are why the output looks right —
|
||||||
|
modern conveniences belong in the workflow, not the output. See
|
||||||
|
[docs/design.md](docs/design.md).
|
||||||
|
|
||||||
## Run
|
## Run
|
||||||
|
|
||||||
|
|
@ -19,6 +21,9 @@ python3 -m http.server 8777 # from this directory
|
||||||
# open http://127.0.0.1:8777
|
# open http://127.0.0.1:8777
|
||||||
```
|
```
|
||||||
|
|
||||||
|
Static files and ES modules — no build step, no dependencies beyond MediaPipe's
|
||||||
|
wasm, which is fetched from a CDN on first use.
|
||||||
|
|
||||||
**Synthetic take** needs no video and exercises everything below detection.
|
**Synthetic take** needs no video and exercises everything below detection.
|
||||||
|
|
||||||
For real footage:
|
For real footage:
|
||||||
|
|
@ -57,11 +62,43 @@ is hand-drawn head plates, which this tool does not yet do.
|
||||||
| Knob | What it does |
|
| Knob | What it does |
|
||||||
| --- | --- |
|
| --- | --- |
|
||||||
| vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. |
|
| vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. |
|
||||||
|
| mouth lead ±f | Shifts the performance tracks earlier against the audio and the head. `[` `]`. |
|
||||||
| contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. |
|
| contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. |
|
||||||
| anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform*. |
|
| anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform*. |
|
||||||
| closed-mouth cut | Aperture below which the mouth interior is emitted as `hidden`. |
|
| closed-mouth cut | Aperture below which the mouth interior is emitted as `hidden`. |
|
||||||
| suggest tolerance | Max head movement before a new plate drawing is required. Affects **Suggest** only. |
|
| suggest tolerance | Max head movement before a new plate drawing is required. Affects **Suggest** only. |
|
||||||
|
|
||||||
|
## Teeth
|
||||||
|
|
||||||
|
MediaPipe has no landmarks inside the lips: the inner ring bounds the cavity and
|
||||||
|
everything within it is just pixels. So teeth come from the image.
|
||||||
|
|
||||||
|
The hazard is vertex correspondence — a traced contour reorders between frames
|
||||||
|
and boils. The way out for a blob specifically is **radial sampling**: march
|
||||||
|
outward from the blob's centroid along N fixed directions and take the last pixel
|
||||||
|
inside. Vertex *k* is then always "the extent in direction *k*", so correspondence
|
||||||
|
holds by construction, the vertex count is fixed, and temporal smoothing is well
|
||||||
|
defined with no reordering possible. It also produces a star-shaped reduction,
|
||||||
|
which is what flat blocks of colour want.
|
||||||
|
|
||||||
|
The **teeth measurement** panel shows exactly what is sampled: region dimmed,
|
||||||
|
kept pixels green, extracted contour amber. Tune against that, not the numbers.
|
||||||
|
|
||||||
|
| Knob | What it does |
|
||||||
|
| --- | --- |
|
||||||
|
| teeth contrast | Gate on the separation between the cavity's dark and bright class means. Otsu always returns *some* threshold, so this is what stops it inventing teeth in a dark mouth. |
|
||||||
|
| cavity erode | Pulls the sampled region in from the lip edge — MediaPipe's inner lip landmarks sit slightly outside the real opening, and lips are bright. |
|
||||||
|
| blob grow/erode | Resizes the found blob. An open pass always runs first to despeckle. |
|
||||||
|
| tongue reject | Drops pixels red relative to their own brightness. Teeth are near-neutral; tongue is not. |
|
||||||
|
| prefer upper | Biases component choice toward the top of the cavity. Area alone picks the tongue when the mouth is wide. |
|
||||||
|
| teeth vertices | Radial sample count. |
|
||||||
|
| teeth avg ±f | Temporal average over the contour. |
|
||||||
|
| teeth dwell | Frames a presence change must persist before it takes effect. |
|
||||||
|
|
||||||
|
**Tongue** as its own part would work the same way, gated on redness instead of
|
||||||
|
brightness and biased low rather than high. Not implemented: it is not visible in
|
||||||
|
the test footage, which reads as a dark cavity with a bright upper-teeth band.
|
||||||
|
|
||||||
## The plate is reference, not art
|
## The plate is reference, not art
|
||||||
|
|
||||||
The plate layer has several representations because its job changes. Cycle with
|
The plate layer has several representations because its job changes. Cycle with
|
||||||
|
|
@ -103,8 +140,16 @@ their own plate drawing. Everything starts kept; delete what you don't want.
|
||||||
**Suggest** runs error-tolerance decimation over head pose as a starting point,
|
**Suggest** runs error-tolerance decimation over head pose as a starting point,
|
||||||
then you hand-correct.
|
then you hand-correct.
|
||||||
|
|
||||||
|
**Mouth lead** is not a correction for a bug. A centred moving average has no
|
||||||
|
phase lag, so smoothing does not delay anything — but it blurs onsets, and the
|
||||||
|
visually salient moment of a mouth opening moves later even though the mean does
|
||||||
|
not. Animators also draw mouth shapes one or two frames ahead of the sound as
|
||||||
|
standard practice. Only the performance tracks shift; the head stays with the
|
||||||
|
audio, since it is the mouth that should anticipate. The lead is baked into the
|
||||||
|
exported take, so the renderer never needs to know about it.
|
||||||
|
|
||||||
`contour avg` is a deliberate, bounded exception to "never smooth the contour" in
|
`contour avg` is a deliberate, bounded exception to "never smooth the contour" in
|
||||||
`../docs/roto-puppet.md`. That rule held while keys were sparse, because sampling
|
`docs/design.md`. That rule held while keys were sparse, because sampling
|
||||||
at velocity minima rejected detector noise for free. With a key on every frame it
|
at velocity minima rejected detector noise for free. With a key on every frame it
|
||||||
does not, so a radius shorter than the shortest articulation worth keeping is
|
does not, so a radius shorter than the shortest articulation worth keeping is
|
||||||
justified — at 12fps, articulation spans 3–6 frames and detector noise is
|
justified — at 12fps, articulation spans 3–6 frames and detector noise is
|
||||||
|
|
@ -117,7 +162,11 @@ chromium --headless --virtual-time-budget=8000 --dump-dom \
|
||||||
http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+'
|
http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+'
|
||||||
```
|
```
|
||||||
|
|
||||||
Or open `selftest.html`. 29 assertions over the stages below detection.
|
Or open `selftest.html`. 41 assertions over the stages below detection, plus a
|
||||||
|
wiring cross-check: every `el('id')` in `app.js` must exist in `index.html`. A
|
||||||
|
knob wired in one but not the other throws during wiring, which aborts the rest
|
||||||
|
of the module and leaves a blank page — a symptom that points nowhere near its
|
||||||
|
cause, and which has happened twice.
|
||||||
|
|
||||||
The ring-simplicity check is the load-bearing one. Because `hold` parts *cut*
|
The ring-simplicity check is the load-bearing one. Because `hold` parts *cut*
|
||||||
between poses instead of interpolating, a ring whose vertex order is wrong
|
between poses instead of interpolating, a ring whose vertex order is wrong
|
||||||
|
|
|
||||||
219
docs/design.md
Normal file
219
docs/design.md
Normal file
|
|
@ -0,0 +1,219 @@
|
||||||
|
# arthur — design
|
||||||
|
|
||||||
|
A small animation suite for turning live-action video into work that reads as
|
||||||
|
hand-authored, in the idiom of *Another World* (Éric Chahi, 1991): flat shapes,
|
||||||
|
a tiny palette, hard edges, motion carried by silhouette.
|
||||||
|
|
||||||
|
## The aesthetic is a representation, not a filter
|
||||||
|
|
||||||
|
Chahi did not process video. He shot reference footage and hand-traced polygons
|
||||||
|
over it in a custom editor; the engine stored and replayed polygon lists, never
|
||||||
|
bitmaps. Three properties follow, and all three are load-bearing:
|
||||||
|
|
||||||
|
1. **Temporal identity.** The same shape, with the same vertex count and vertex
|
||||||
|
order, *edited* across frames. A shape re-detected independently each frame
|
||||||
|
produces a new contour every frame. That boils, and boiling reads as "filter"
|
||||||
|
within half a second no matter how good the individual shapes are.
|
||||||
|
2. **Stylised timing.** Poses held, and for most parts a hard cut between them
|
||||||
|
rather than an interpolation. A distinct pose on every frame of 24fps footage
|
||||||
|
looks like video even when every pose is a polygon.
|
||||||
|
3. **Authored colour.** A small fixed ramp with two or three tones per part,
|
||||||
|
chosen by a person. Sampling colour from the source produces a pixel-art
|
||||||
|
filter immediately and irrecoverably.
|
||||||
|
|
||||||
|
None of those are computer-vision problems. That is the whole argument about
|
||||||
|
where CV belongs.
|
||||||
|
|
||||||
|
## The constraint is the point
|
||||||
|
|
||||||
|
320×200, indexed palette, flat fills, no antialiasing. That is inherited from
|
||||||
|
Animator Pro, where this work started, but it is not an accident to be
|
||||||
|
modernised away — it is why the output looks right. The rasteriser writes palette
|
||||||
|
indices into a byte buffer and expands to RGBA only at the end, precisely so no
|
||||||
|
canvas antialiasing can soften an edge.
|
||||||
|
|
||||||
|
Modern conveniences belong in the *workflow* — instant feedback, audio, real
|
||||||
|
undo, scrubbing, layers. Not in the output.
|
||||||
|
|
||||||
|
## CV tracks; it does not draw
|
||||||
|
|
||||||
|
- **Its job.** Where the rigid features of the face are, what the head's
|
||||||
|
transform is, where the lip contour runs, which frame is a motion extreme.
|
||||||
|
- **Its non-job.** Deciding what the shapes are. A face mesh offers 468 points;
|
||||||
|
an animator's mouth is eight. **That decimation ratio is the style.** It is a
|
||||||
|
decision encoded in the tool, not a measurement extracted from footage.
|
||||||
|
|
||||||
|
Every number crossing from analysis into rendering is a *parameter* or a
|
||||||
|
*correspondence*. Every number determining how something looks comes from the
|
||||||
|
artist or a fixed authored table.
|
||||||
|
|
||||||
|
## Two kinds of part
|
||||||
|
|
||||||
|
| Kind | Source | Vocabulary | Interp |
|
||||||
|
| --- | --- | --- | --- |
|
||||||
|
| **Plate** — head, hair, body | Hand-drawn | Closed: a few drawings per character | hold |
|
||||||
|
| **Feature** — mouth, lids | Rotoscoped from landmarks | Open: derived from this take | hold |
|
||||||
|
| **Interior** — mouth interior, teeth | Image content within a feature | Open | hold |
|
||||||
|
| **Primitive** — iris | Landmark centroid as a disc | Quantised | hold |
|
||||||
|
|
||||||
|
The asymmetry is deliberate, and it is the opposite choice in each case.
|
||||||
|
|
||||||
|
A **closed vocabulary is wrong for the mouth.** Pre-authored mouth shapes are
|
||||||
|
never *this performance's* shapes, and that genericness is what the project
|
||||||
|
exists to avoid. So the mouth gets an open vocabulary derived from footage, and
|
||||||
|
the stylisation applies to its **timing** instead.
|
||||||
|
|
||||||
|
A **closed vocabulary is right for the head.** The head carries structure, not
|
||||||
|
performance. A handful of hand-drawn angles is better, because you drew them —
|
||||||
|
and it means the head never needs segmentation, contour tracking, or boil
|
||||||
|
avoidance.
|
||||||
|
|
||||||
|
## Two kinds of sparseness
|
||||||
|
|
||||||
|
Conflating these was the original design error. **Aesthetic** sparseness is set
|
||||||
|
by the extraction rate: pick 12fps and the timing is already chosen. **Labour**
|
||||||
|
sparseness is a human drawing each one, and it binds only on the plate.
|
||||||
|
|
||||||
|
So the mouth keeps **every** frame — it is traced, and therefore free. In limited
|
||||||
|
animation lip sync is routinely the densest element, on 1s, while heads hold on
|
||||||
|
2s and 3s.
|
||||||
|
|
||||||
|
The other half is frame removal: which frames need their own plate drawing.
|
||||||
|
Everything starts kept; delete what you don't want. Automatic suggestion runs
|
||||||
|
error-tolerance decimation over head pose as a starting point, then you
|
||||||
|
hand-correct.
|
||||||
|
|
||||||
|
## Stabilisation
|
||||||
|
|
||||||
|
A rotoscoped mouth is only reusable if expressed independently of where the head
|
||||||
|
was. That inverts the obvious composition:
|
||||||
|
|
||||||
|
```text
|
||||||
|
mouth_local(t) = anchor(t)⁻¹ · mouth_world(t)
|
||||||
|
```
|
||||||
|
|
||||||
|
- **Rigid landmarks only** for the fit: eye corners, nose bridge, nose tip.
|
||||||
|
Including a feature that moves bleeds performance into the stabilisation.
|
||||||
|
- **Similarity, not affine or homography.** The extra degrees of freedom absorb
|
||||||
|
head rotation as distortion and smear it into the mouth. Four DOF removes
|
||||||
|
exactly translation, roll and depth scale, leaving yaw and pitch as a
|
||||||
|
measurable residual.
|
||||||
|
- **Reference is the mean** configuration over the shot, not frame zero.
|
||||||
|
- **Smooth the transform, never the contour** — with one bounded exception,
|
||||||
|
below.
|
||||||
|
- Landmarks convert to an **isotropic** space first (unit = one image height).
|
||||||
|
MediaPipe normalises x by width and y by height, so its space is stretched; a
|
||||||
|
"similarity" fitted there is not one.
|
||||||
|
|
||||||
|
Out-of-plane rotation cannot be removed by any 2D transform. The answer is not a
|
||||||
|
better transform: it is to draw the head at each angle and let each drawing
|
||||||
|
declare where its mouth sits. Foreshortening becomes authored metadata.
|
||||||
|
|
||||||
|
## Hold versus interpolate
|
||||||
|
|
||||||
|
Another World had no inbetweening engine: polygon sets played back frame by
|
||||||
|
frame at a low rate. For parts that should read as hand-animated snaps — mouths
|
||||||
|
above all — **cut between keys, do not interpolate.** A tweened mouth is rubbery
|
||||||
|
and reads as puppet software immediately.
|
||||||
|
|
||||||
|
**Fixed topology is non-negotiable for cut parts.** Vertex *meanings* must be
|
||||||
|
stable across every key, so a closed mouth and a wide-open mouth are the same
|
||||||
|
polygon at different positions. Interpolation partially hides ordering drift;
|
||||||
|
cutting exposes it completely, as static.
|
||||||
|
|
||||||
|
## Deriving shapes from pixels without boiling
|
||||||
|
|
||||||
|
Some things have no landmarks — teeth, tongue, anything inside the lips. The
|
||||||
|
hazard is vertex correspondence: a traced contour reorders between frames.
|
||||||
|
|
||||||
|
Two escapes, both used:
|
||||||
|
|
||||||
|
- **Extract a scalar, not a shape**, where the shape can be derived from
|
||||||
|
geometry you already trust.
|
||||||
|
- **Radial sampling** where a real contour is needed. March outward from the
|
||||||
|
blob's centroid along N fixed directions; vertex *k* is then always "the extent
|
||||||
|
in direction *k*". Correspondence holds by construction, the count is fixed,
|
||||||
|
temporal smoothing is well defined, and the star-shaped result suits flat
|
||||||
|
colour.
|
||||||
|
|
||||||
|
## The bounded smoothing exception
|
||||||
|
|
||||||
|
*Smooth the transform, never the contour* held while keys were sparse: sampling
|
||||||
|
at velocity minima rejected per-frame detector noise for free. With a key on
|
||||||
|
every frame it does not, so a short contour average is justified — the window
|
||||||
|
must stay **shorter than the shortest articulation worth keeping**. At 12fps,
|
||||||
|
mouth movement spans 3–6 frames and detector noise is per-frame, so ±1 separates
|
||||||
|
them and ±3 starts eating speech. At 24fps, double it.
|
||||||
|
|
||||||
|
## Performance tracks may lead the clock
|
||||||
|
|
||||||
|
A centred moving average has no phase lag, but it blurs onsets, so the visually
|
||||||
|
salient moment of a mouth opening moves later even though the mean does not.
|
||||||
|
Animators also draw mouth shapes a frame or two ahead of the sound as standard
|
||||||
|
practice. Only performance tracks shift; the head stays with the audio, because
|
||||||
|
it is the mouth that should anticipate.
|
||||||
|
|
||||||
|
## Colour discipline
|
||||||
|
|
||||||
|
- A small ramp. Another World ran 16 colours at 320×200.
|
||||||
|
- Two or three tones per part: base, shadow, occasionally a rim.
|
||||||
|
- No dithering, no antialiasing anywhere.
|
||||||
|
- Never sample colour from the source. Analysis may report *which* tone a region
|
||||||
|
should be; it must never report an RGB value.
|
||||||
|
- Plate art and generated parts share one palette, authored together.
|
||||||
|
|
||||||
|
## Three artifacts, not two
|
||||||
|
|
||||||
|
| Artifact | Cost | Regenerated when |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| Dense track — landmarks and fitted anchor per frame | Minutes, once per shot | Re-shoot, or a model change |
|
||||||
|
| Take — parts, keys, kept frames | Instant | Every knob change |
|
||||||
|
| Render | — | Continuous |
|
||||||
|
|
||||||
|
Keep the dense track: re-keying is then a regenerate, never a re-trace. And
|
||||||
|
because the take is disposable, **hand corrections must not live in it** — they
|
||||||
|
belong in an override layer keyed by `(part, frame)`, applied on top at render
|
||||||
|
time.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
The organising idea is an **indexed raster core**, a **part/track model**, and
|
||||||
|
**pluggable sources** feeding it. Sources so far are rotoscope and image-derived
|
||||||
|
interior; hand-drawn and procedural are the obvious next ones. Everything else —
|
||||||
|
stabilisation, key selection, frame removal, palette — operates on the track
|
||||||
|
model and is source-agnostic.
|
||||||
|
|
||||||
|
Current modules:
|
||||||
|
|
||||||
|
| Module | Role |
|
||||||
|
| --- | --- |
|
||||||
|
| `landmarks.js` | Index tables. Ring arrays are ordered traversals: slot position *is* vertex identity. |
|
||||||
|
| `mathutil.js` | Similarity fit, Procrustes mean, temporal smoothing. |
|
||||||
|
| `pipeline.js` | Stabilise → subsample → key-select → frame-removal. |
|
||||||
|
| `interior.js` | Teeth from image content: Otsu, morphology, components, radial contour. |
|
||||||
|
| `underlay.js` | Registered photo reference and palette posterisation. |
|
||||||
|
| `raster.js` | Indexed scanline fill. No antialiasing, by construction. |
|
||||||
|
| `take.js` | Take-file writer. |
|
||||||
|
| `synth.js` | Synthetic landmarks, so everything below detection is testable without a video. |
|
||||||
|
| `selftest.js` | Assertions, including source-level wiring checks. |
|
||||||
|
|
||||||
|
## Not built yet
|
||||||
|
|
||||||
|
- **A paint surface.** The plates have nowhere to be drawn. This is the largest
|
||||||
|
gap between "tool" and "suite": a pixel paint canvas with onion skin, palette
|
||||||
|
constraint, and the registered underlay behind it.
|
||||||
|
- Eyes, irises, brows as parts.
|
||||||
|
- Plate libraries with per-plate mouth slots.
|
||||||
|
- Real performer→character calibration (currently identity).
|
||||||
|
- The override layer.
|
||||||
|
- Export beyond the take file — FLI/FLC would connect back to the lineage and is
|
||||||
|
a simple format.
|
||||||
|
|
||||||
|
## Lineage
|
||||||
|
|
||||||
|
This began as an analysis front-end for Animator Pro, whose render path was the
|
||||||
|
intended destination. `Animator-Pro-SDL/docs/roto-puppet.md` documents that
|
||||||
|
design and what the Poco/FLX machinery can and cannot do. The reason for leaving
|
||||||
|
is in that document's own rule — *never make a timing decision that requires a
|
||||||
|
full render to evaluate* — which, once honoured, left the host with nothing to do
|
||||||
|
but write the file.
|
||||||
6
gl.html
6
gl.html
|
|
@ -1,6 +0,0 @@
|
||||||
<!doctype html><html><body><pre id=o>?</pre><script>
|
|
||||||
const c=document.createElement('canvas');
|
|
||||||
const g=c.getContext('webgl2')||c.getContext('webgl');
|
|
||||||
document.getElementById('o').textContent = g ? 'WEBGL OK '+g.getParameter(g.VERSION) : 'WEBGL UNAVAILABLE';
|
|
||||||
document.title = g ? 'GLOK' : 'GLNO';
|
|
||||||
</script></body></html>
|
|
||||||
37
index.html
37
index.html
|
|
@ -3,7 +3,7 @@
|
||||||
<head>
|
<head>
|
||||||
<meta charset="utf-8">
|
<meta charset="utf-8">
|
||||||
<meta name="viewport" content="width=device-width, initial-scale=1">
|
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||||||
<title>roto — take builder</title>
|
<title>arthur</title>
|
||||||
<style>
|
<style>
|
||||||
:root { --bg:#0b0d13; --panel:#12151f; --line:#232836; --fg:#e8eaf0;
|
:root { --bg:#0b0d13; --panel:#12151f; --line:#232836; --fg:#e8eaf0;
|
||||||
--dim:#8891a5; --accent:#fbbf24; --ok:#4ade80; --err:#f87171; }
|
--dim:#8891a5; --accent:#fbbf24; --ok:#4ade80; --err:#f87171; }
|
||||||
|
|
@ -58,7 +58,7 @@
|
||||||
</head>
|
</head>
|
||||||
<body>
|
<body>
|
||||||
<header>
|
<header>
|
||||||
<h1>roto <span>— frame removal + take builder</span></h1>
|
<h1>arthur <span>— rotoscope + frame removal</span></h1>
|
||||||
<button id="btn-synth">Synthetic</button>
|
<button id="btn-synth">Synthetic</button>
|
||||||
<input type="text" id="framedir" value="frames" size="7" title="frame directory">
|
<input type="text" id="framedir" value="frames" size="7" title="frame directory">
|
||||||
<button id="btn-frames">Load frames</button>
|
<button id="btn-frames">Load frames</button>
|
||||||
|
|
@ -91,7 +91,7 @@
|
||||||
<div class="panel">
|
<div class="panel">
|
||||||
<h2>stabilised (head-local)</h2>
|
<h2>stabilised (head-local)</h2>
|
||||||
<canvas id="cv-stab"></canvas>
|
<canvas id="cv-stab"></canvas>
|
||||||
<div class="legend">should sit still except the mouth · grey = held plate outline</div>
|
<div class="legend">should sit still except the mouth · grey = held plate outline<br>dark green ghost = unshifted mouth when lead ≠ 0</div>
|
||||||
</div>
|
</div>
|
||||||
<div class="panel">
|
<div class="panel">
|
||||||
<h2>flat render — 320×200 indexed</h2>
|
<h2>flat render — 320×200 indexed</h2>
|
||||||
|
|
@ -109,6 +109,7 @@
|
||||||
<input type="range" id="scrub" min="0" max="0" value="0" style="width:100%;margin-top:10px">
|
<input type="range" id="scrub" min="0" max="0" value="0" style="width:100%;margin-top:10px">
|
||||||
<div class="legend">
|
<div class="legend">
|
||||||
<kbd>←</kbd> <kbd>→</kbd> step · <kbd>X</kbd> delete · <kbd>K</kbd> keep · <kbd>B</kbd> background ·
|
<kbd>←</kbd> <kbd>→</kbd> step · <kbd>X</kbd> delete · <kbd>K</kbd> keep · <kbd>B</kbd> background ·
|
||||||
|
<kbd>[</kbd> <kbd>]</kbd> lead ·
|
||||||
click to select · double-click or shift-click to toggle ·
|
click to select · double-click or shift-click to toggle ·
|
||||||
the mouth keeps <b>every</b> frame regardless
|
the mouth keeps <b>every</b> frame regardless
|
||||||
</div>
|
</div>
|
||||||
|
|
@ -119,19 +120,47 @@
|
||||||
<div class="panel" style="flex:1 1 400px">
|
<div class="panel" style="flex:1 1 400px">
|
||||||
<h2>knobs</h2>
|
<h2>knobs</h2>
|
||||||
<label class="ctl"><span>vertices</span><input type="range" id="verts" min="4" max="16" step="2" value="8"><output id="vertsv"></output></label>
|
<label class="ctl"><span>vertices</span><input type="range" id="verts" min="4" max="16" step="2" value="8"><output id="vertsv"></output></label>
|
||||||
|
<label class="ctl"><span>mouth lead ±f</span><input type="range" id="lead" min="-6" max="6" value="0"><output id="leadv"></output></label>
|
||||||
<label class="ctl"><span>contour avg ±f</span><input type="range" id="contourSmooth" min="0" max="4" value="1"><output id="contourSmoothv"></output></label>
|
<label class="ctl"><span>contour avg ±f</span><input type="range" id="contourSmooth" min="0" max="4" value="1"><output id="contourSmoothv"></output></label>
|
||||||
<label class="ctl"><span>anchor avg ±f</span><input type="range" id="smoothWin" min="0" max="8" value="2"><output id="smoothWinv"></output></label>
|
<label class="ctl"><span>anchor avg ±f</span><input type="range" id="smoothWin" min="0" max="8" value="2"><output id="smoothWinv"></output></label>
|
||||||
<label class="ctl"><span>closed-mouth cut</span><input type="range" id="apertureThresh" min="0" max="400" value="120"><output id="apertureThreshv"></output></label>
|
<label class="ctl"><span>closed-mouth cut</span><input type="range" id="apertureThresh" min="0" max="400" value="120"><output id="apertureThreshv"></output></label>
|
||||||
|
<label class="ctl"><span>teeth contrast</span><input type="range" id="teethOn" min="1" max="60" value="16"><output id="teethOnv"></output></label>
|
||||||
|
<label class="ctl"><span>cavity erode</span><input type="range" id="teethErode" min="0" max="45" value="18"><output id="teethErodev"></output></label>
|
||||||
|
<label class="ctl"><span>blob grow/erode</span><input type="range" id="blobGrow" min="-4" max="4" value="0"><output id="blobGrowv"></output></label>
|
||||||
|
<label class="ctl"><span>tongue reject</span><input type="range" id="tongueReject" min="2" max="40" value="18"><output id="tongueRejectv"></output></label>
|
||||||
|
<label class="ctl"><span>prefer upper</span><input type="range" id="topBias" min="0" max="120" value="60"><output id="topBiasv"></output></label>
|
||||||
|
<label class="ctl"><span>teeth vertices</span><input type="range" id="teethVerts" min="5" max="20" value="10"><output id="teethVertsv"></output></label>
|
||||||
|
<label class="ctl"><span>teeth avg ±f</span><input type="range" id="teethSmooth" min="0" max="4" value="1"><output id="teethSmoothv"></output></label>
|
||||||
|
<label class="ctl"><span>teeth dwell</span><input type="range" id="teethDwell" min="0" max="6" value="1"><output id="teethDwellv"></output></label>
|
||||||
<label class="ctl"><span>suggest tolerance</span><input type="range" id="tol" min="2" max="60" value="14"><output id="tolv"></output></label>
|
<label class="ctl"><span>suggest tolerance</span><input type="range" id="tol" min="2" max="60" value="14"><output id="tolv"></output></label>
|
||||||
<div class="legend">
|
<div class="legend">
|
||||||
|
<b>mouth lead</b> shifts the performance tracks earlier (positive) against
|
||||||
|
the audio and the head. Averaging has no phase lag but it blurs onsets, so
|
||||||
|
an opening reads later than it is; animators also draw mouths a frame or
|
||||||
|
two ahead of the sound as standard practice. <kbd>[</kbd> <kbd>]</kbd>.<br>
|
||||||
<b>contour avg</b> 0 = off, 1 = ±1 frame. Removes per-frame landmark
|
<b>contour avg</b> 0 = off, 1 = ±1 frame. Removes per-frame landmark
|
||||||
jitter. Push past 2 and it starts eating articulation.<br>
|
jitter. Push past 2 and it starts eating articulation.<br>
|
||||||
<b>anchor avg</b> smooths the head transform only — never the contour.<br>
|
<b>anchor avg</b> smooths the head transform only — never the contour.<br>
|
||||||
|
<b>teeth contrast</b> gates on how far apart the cavity's dark and bright
|
||||||
|
halves are — Otsu always returns <i>some</i> threshold, so this is what
|
||||||
|
stops it inventing teeth in a dark mouth.
|
||||||
|
<b>cavity erode</b> pulls the sampled region in from the lip edge;
|
||||||
|
<b>blob grow/erode</b> resizes the found blob itself.
|
||||||
|
<b>tongue reject</b> drops pixels that are red relative to their own
|
||||||
|
brightness; <b>prefer upper</b> biases component choice toward the top of
|
||||||
|
the cavity, where teeth are and the tongue is not.
|
||||||
|
<b>dwell</b> is how many frames a presence change must persist.<br>
|
||||||
<b>suggest tolerance</b> only affects the Suggest button: max head movement
|
<b>suggest tolerance</b> only affects the Suggest button: max head movement
|
||||||
allowed before a new drawing is required.
|
allowed before a new drawing is required.
|
||||||
</div>
|
</div>
|
||||||
</div>
|
</div>
|
||||||
<div class="panel" style="flex:1 1 280px">
|
<div class="panel" style="flex:0 1 190px">
|
||||||
|
<h2>teeth measurement</h2>
|
||||||
|
<div id="cv-teeth"></div>
|
||||||
|
<div class="legend" id="teethinfo"></div>
|
||||||
|
<div class="legend">green = kept pixels · amber = extracted contour</div>
|
||||||
|
</div>
|
||||||
|
<div class="panel" style="flex:1 1 240px">
|
||||||
<h2>palette</h2>
|
<h2>palette</h2>
|
||||||
<div id="palette"></div>
|
<div id="palette"></div>
|
||||||
<div class="legend" style="margin-top:12px">
|
<div class="legend" style="margin-top:12px">
|
||||||
|
|
|
||||||
217
js/app.js
217
js/app.js
|
|
@ -1,8 +1,10 @@
|
||||||
import { FaceLandmarker, FilesetResolver } from 'https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@1.0.1/vision_bundle.mjs';
|
import { FaceLandmarker, FilesetResolver } from 'https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@1.0.1/vision_bundle.mjs';
|
||||||
import { LIPS_OUTER, LIPS_INNER, FACE_OVAL } from './landmarks.js';
|
import { LIPS_OUTER, LIPS_INNER, FACE_OVAL } from './landmarks.js';
|
||||||
import { stabilize, toRasterRing, smoothContours, suggestPlateFrames, heldFrame } from './pipeline.js';
|
import { stabilize, toRasterRing, smoothContours, suggestPlateFrames, heldFrame, shiftIndex } from './pipeline.js';
|
||||||
import { IndexedRaster } from './raster.js';
|
import { IndexedRaster } from './raster.js';
|
||||||
import { drawRegistered, posterizeInto } from './underlay.js';
|
import { drawRegistered, posterizeInto } from './underlay.js';
|
||||||
|
import { extractTeeth } from './interior.js';
|
||||||
|
import { applySim } from './mathutil.js';
|
||||||
import { writeTake } from './take.js';
|
import { writeTake } from './take.js';
|
||||||
import { synthDense } from './synth.js';
|
import { synthDense } from './synth.js';
|
||||||
|
|
||||||
|
|
@ -13,9 +15,9 @@ const PALETTE = [
|
||||||
{ name: 'skin_base', hex: '#b07a5a' },
|
{ name: 'skin_base', hex: '#b07a5a' },
|
||||||
{ name: 'skin_dark', hex: '#7a4f3a' },
|
{ name: 'skin_dark', hex: '#7a4f3a' },
|
||||||
{ name: 'mouth_dark', hex: '#24161a' },
|
{ name: 'mouth_dark', hex: '#24161a' },
|
||||||
{ name: 'skin_lite', hex: '#d9a884' },
|
{ name: 'teeth', hex: '#d9cfc2' },
|
||||||
];
|
];
|
||||||
const IDX = { bg: 0, base: 1, dark: 2, mouth: 3, lite: 4 };
|
const IDX = { bg: 0, base: 1, dark: 2, mouth: 3, teeth: 4 };
|
||||||
|
|
||||||
const state = {
|
const state = {
|
||||||
dense: null, images: [], stab: null, xform: null,
|
dense: null, images: [], stab: null, xform: null,
|
||||||
|
|
@ -23,11 +25,30 @@ const state = {
|
||||||
keep: new Set(), // frames that get their own plate drawing
|
keep: new Set(), // frames that get their own plate drawing
|
||||||
frame: 0, playing: false, faceBox: null,
|
frame: 0, playing: false, faceBox: null,
|
||||||
fps: 12, audio: null, // fps comes from manifest.json, never guessed
|
fps: 12, audio: null, // fps comes from manifest.json, never guessed
|
||||||
|
aspect: 1, // imgW/imgH; converts MediaPipe's anisotropic space
|
||||||
|
lead: 0, // performance-track offset in frames
|
||||||
|
interior: null, // per-frame teeth measurement from image content
|
||||||
|
teeth: null, // resolved per-frame {show, t} after knobs
|
||||||
};
|
};
|
||||||
|
|
||||||
const el = (id) => document.getElementById(id);
|
const el = (id) => {
|
||||||
|
const n = document.getElementById(id);
|
||||||
|
// A knob present in the code but missing from the markup used to throw during
|
||||||
|
// wiring and leave a blank page with nothing in the console worth reading.
|
||||||
|
if (!n) throw new Error(`missing element #${id} — knob wired in app.js but not in index.html`);
|
||||||
|
return n;
|
||||||
|
};
|
||||||
const opts = () => ({
|
const opts = () => ({
|
||||||
verts: +el('verts').value,
|
verts: +el('verts').value,
|
||||||
|
lead: +el('lead').value,
|
||||||
|
teethOn: +el('teethOn').value / 100, // minimum Otsu class separation
|
||||||
|
teethDwell: +el('teethDwell').value,
|
||||||
|
teethSmooth: +el('teethSmooth').value,
|
||||||
|
cavityErode: +el('teethErode').value / 100,
|
||||||
|
tongueReject: +el('tongueReject').value / 100,
|
||||||
|
blobGrow: +el('blobGrow').value,
|
||||||
|
topBias: +el('topBias').value / 100,
|
||||||
|
teethVerts: +el('teethVerts').value,
|
||||||
smoothWin: +el('smoothWin').value,
|
smoothWin: +el('smoothWin').value,
|
||||||
contourSmooth: +el('contourSmooth').value,
|
contourSmooth: +el('contourSmooth').value,
|
||||||
apertureThresh: +el('apertureThresh').value / 1000,
|
apertureThresh: +el('apertureThresh').value / 1000,
|
||||||
|
|
@ -115,6 +136,19 @@ async function detectAll(images) {
|
||||||
return { dense, missing };
|
return { dense, missing };
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Interior measurement is a function of pixels alone, so it runs once with
|
||||||
|
// detection and the knobs re-resolve it instantly afterwards.
|
||||||
|
function measureAll(images, dense, o) {
|
||||||
|
const ctx = document.createElement('canvas').getContext('2d', { willReadFrequently: true });
|
||||||
|
return dense.map((lm, i) =>
|
||||||
|
extractTeeth(images[i], LIPS_INNER.map((k) => lm[k]), ctx, o));
|
||||||
|
}
|
||||||
|
|
||||||
|
// Extraction keys on every knob that changes the pixels examined, so the cache
|
||||||
|
// is keyed on exactly those and a change to anything else stays instant.
|
||||||
|
const extractKey = (o) =>
|
||||||
|
[o.cavityErode, o.tongueReject, o.blobGrow, o.topBias, o.teethVerts].join('|');
|
||||||
|
|
||||||
/* ---------- build ---------- */
|
/* ---------- build ---------- */
|
||||||
|
|
||||||
function makeXform(stab, neutral) {
|
function makeXform(stab, neutral) {
|
||||||
|
|
@ -132,9 +166,10 @@ function makeXform(stab, neutral) {
|
||||||
function rebuild(resetKeep) {
|
function rebuild(resetKeep) {
|
||||||
if (!state.dense) return;
|
if (!state.dense) return;
|
||||||
const o = opts();
|
const o = opts();
|
||||||
|
state.lead = o.lead;
|
||||||
const N = state.dense.length;
|
const N = state.dense.length;
|
||||||
|
|
||||||
state.stab = stabilize(state.dense, o.smoothWin);
|
state.stab = stabilize(state.dense, o.smoothWin, state.aspect);
|
||||||
|
|
||||||
const ap = state.stab.aperture;
|
const ap = state.stab.aperture;
|
||||||
const head = Math.max(1, Math.floor(N / 4));
|
const head = Math.max(1, Math.floor(N / 4));
|
||||||
|
|
@ -153,6 +188,12 @@ function rebuild(resetKeep) {
|
||||||
const apMax = Math.max(...ap);
|
const apMax = Math.max(...ap);
|
||||||
state.hidden = ap.map((v) => v / apMax < o.apertureThresh);
|
state.hidden = ap.map((v) => v / apMax < o.apertureThresh);
|
||||||
|
|
||||||
|
if (state.images.length && state.extractKey !== extractKey(o)) {
|
||||||
|
state.interior = measureAll(state.images, state.dense, o);
|
||||||
|
state.extractKey = extractKey(o);
|
||||||
|
}
|
||||||
|
state.teeth = resolveTeeth(o);
|
||||||
|
|
||||||
// Plate outline per frame, so a kept frame shows its own head shape.
|
// Plate outline per frame, so a kept frame shows its own head shape.
|
||||||
state.plates = state.stab.oval.map((r) => r.map(state.xform));
|
state.plates = state.stab.oval.map((r) => r.map(state.xform));
|
||||||
|
|
||||||
|
|
@ -179,6 +220,53 @@ function faceBoxes() {
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Presence gets hysteresis and a minimum dwell, the same treatment plate
|
||||||
|
// selection gets: a teeth block that blinks on and off for single frames is
|
||||||
|
// worse than one that is simply absent. Appearing needs a clear signal, staying
|
||||||
|
// needs only a weak one.
|
||||||
|
function resolveTeeth(o) {
|
||||||
|
const N = state.dense.length;
|
||||||
|
if (!state.interior) return new Array(N).fill({ show: false, pts: null });
|
||||||
|
|
||||||
|
const raw = state.interior.map((m, f) =>
|
||||||
|
(state.hidden[f] || !m.contour ? 0 : m.contrast));
|
||||||
|
const on = o.teethOn, off = o.teethOn * 0.7;
|
||||||
|
const shown = new Array(N).fill(false);
|
||||||
|
let live = false, since = 0;
|
||||||
|
for (let f = 0; f < N; f++) {
|
||||||
|
const want = live ? raw[f] > off : raw[f] > on;
|
||||||
|
if (want !== live && since >= o.teethDwell) { live = want; since = 0; }
|
||||||
|
else since++;
|
||||||
|
shown[f] = live && !state.hidden[f] && !!state.interior[f].contour;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Into raster space through the same chain the lips take, including the
|
||||||
|
// isotropic aspect conversion - a contour in MediaPipe's normalised space is
|
||||||
|
// in the same stretched coordinates the landmarks are.
|
||||||
|
const toRaster = (pts, f) => {
|
||||||
|
const tf = state.stab.transforms[f];
|
||||||
|
return pts.map((p) => state.xform(applySim(tf, { x: p.x * state.aspect, y: p.y })));
|
||||||
|
};
|
||||||
|
const rast = state.interior.map((m, f) => (m.contour ? toRaster(m.contour, f) : null));
|
||||||
|
|
||||||
|
// Radial sampling makes vertex k mean the same direction on every frame, so
|
||||||
|
// smoothing across time is well defined and cannot reorder anything.
|
||||||
|
const sm = rast.map((pts, f) => {
|
||||||
|
if (!pts || !shown[f]) return pts;
|
||||||
|
const acc = pts.map(() => ({ x: 0, y: 0 }));
|
||||||
|
let c = 0;
|
||||||
|
for (let j = f - o.teethSmooth; j <= f + o.teethSmooth; j++) {
|
||||||
|
const k = Math.min(N - 1, Math.max(0, j));
|
||||||
|
if (!shown[k] || !rast[k] || rast[k].length !== pts.length) continue;
|
||||||
|
for (let v = 0; v < pts.length; v++) { acc[v].x += rast[k][v].x; acc[v].y += rast[k][v].y; }
|
||||||
|
c++;
|
||||||
|
}
|
||||||
|
return c ? acc.map((p) => ({ x: p.x / c, y: p.y / c })) : pts;
|
||||||
|
});
|
||||||
|
|
||||||
|
return shown.map((show, f) => ({ show, pts: sm[f] }));
|
||||||
|
}
|
||||||
|
|
||||||
/* ---------- render ---------- */
|
/* ---------- render ---------- */
|
||||||
|
|
||||||
// The plate layer has several representations because its job changes: a flat
|
// The plate layer has several representations because its job changes: a flat
|
||||||
|
|
@ -197,8 +285,13 @@ function renderFrame(f, mode = plateMode()) {
|
||||||
if (mode === 'oval' || mode === 'oval+photo') r.fillPoly(state.plates[pf], IDX.base);
|
if (mode === 'oval' || mode === 'oval+photo') r.fillPoly(state.plates[pf], IDX.base);
|
||||||
}
|
}
|
||||||
|
|
||||||
r.fillPoly(state.outer[f], IDX.dark); // mouth keeps every frame
|
const mf = leadIndex(f); // performance frame, possibly ahead
|
||||||
if (!state.hidden[f]) r.fillPoly(state.inner[f], IDX.mouth);
|
r.fillPoly(state.outer[mf], IDX.dark); // mouth keeps every frame
|
||||||
|
if (!state.hidden[mf]) {
|
||||||
|
r.fillPoly(state.inner[mf], IDX.mouth);
|
||||||
|
const te = state.teeth[mf];
|
||||||
|
if (te.show && te.pts && te.pts.length >= 3) r.fillPoly(te.pts, IDX.teeth);
|
||||||
|
}
|
||||||
return r;
|
return r;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
@ -243,6 +336,22 @@ function compositeRender(canvas, f, zoom) {
|
||||||
|
|
||||||
const keptSorted = () => [...state.keep].sort((a, b) => a - b);
|
const keptSorted = () => [...state.keep].sort((a, b) => a - b);
|
||||||
|
|
||||||
|
// Performance tracks can lead the clock.
|
||||||
|
//
|
||||||
|
// A centred moving average has no phase lag, so smoothing does not literally
|
||||||
|
// delay anything - but it blurs onsets, and the visually salient moment of a
|
||||||
|
// mouth opening moves later even though the mean does not. Animators also draw
|
||||||
|
// mouth shapes one or two frames ahead of the sound as a matter of course, so
|
||||||
|
// this is the normal control rather than a correction.
|
||||||
|
//
|
||||||
|
// Positive lead = the mouth arrives earlier. Only performance parts shift; the
|
||||||
|
// head stays with the audio, because it is the mouth that should anticipate.
|
||||||
|
function leadIndex(f) {
|
||||||
|
// Reads a cached scalar, not opts(): this runs once per strip thumbnail, and
|
||||||
|
// calling opts() here meant ~14 DOM reads x 74 frames on every redraw.
|
||||||
|
return shiftIndex(f, state.lead, state.dense.length);
|
||||||
|
}
|
||||||
|
|
||||||
function blit(canvas, raster, zoom) {
|
function blit(canvas, raster, zoom) {
|
||||||
canvas.width = RW * zoom; canvas.height = RH * zoom;
|
canvas.width = RW * zoom; canvas.height = RH * zoom;
|
||||||
canvas.getContext('2d').putImageData(raster.toImageData(PALETTE.map((p) => p.hex), zoom), 0, 0);
|
canvas.getContext('2d').putImageData(raster.toImageData(PALETTE.map((p) => p.hex), zoom), 0, 0);
|
||||||
|
|
@ -258,8 +367,12 @@ function drawAll() {
|
||||||
function drawReadout() {
|
function drawReadout() {
|
||||||
const kept = keptSorted();
|
const kept = keptSorted();
|
||||||
const runs = kept.map((k, i) => (i + 1 < kept.length ? kept[i + 1] : state.dense.length) - k);
|
const runs = kept.map((k, i) => (i + 1 < kept.length ? kept[i + 1] : state.dense.length) - k);
|
||||||
|
const lead = state.lead;
|
||||||
|
const teethFrames = state.teeth ? state.teeth.filter((t) => t.show).length : 0;
|
||||||
el('readout').textContent =
|
el('readout').textContent =
|
||||||
`${state.dense.length} frames → ${kept.length} drawings · ` +
|
`${state.dense.length} frames → ${kept.length} drawings · ` +
|
||||||
|
`teeth on ${teethFrames}f · ` +
|
||||||
|
(lead ? `mouth leads ${lead}f (${(lead / state.fps * 1000).toFixed(0)}ms) · ` : '') +
|
||||||
`holds ${Math.min(...runs)}–${Math.max(...runs)} frames · ` +
|
`holds ${Math.min(...runs)}–${Math.max(...runs)} frames · ` +
|
||||||
`neutral f${state.neutral} · residual ` +
|
`neutral f${state.neutral} · residual ` +
|
||||||
`${(state.stab.residual.reduce((a, b) => a + b, 0) / state.dense.length).toFixed(4)}`;
|
`${(state.stab.residual.reduce((a, b) => a + b, 0) / state.dense.length).toFixed(4)}`;
|
||||||
|
|
@ -268,8 +381,13 @@ function drawReadout() {
|
||||||
function drawPanes() {
|
function drawPanes() {
|
||||||
const f = state.frame, kept = keptSorted();
|
const f = state.frame, kept = keptSorted();
|
||||||
const pf = heldFrame(kept, f);
|
const pf = heldFrame(kept, f);
|
||||||
|
// The mouth frame is always shown, not only when shifted, so the number can be
|
||||||
|
// watched diverging from f rather than taken on trust.
|
||||||
|
const lead = state.lead;
|
||||||
el('framelabel').textContent =
|
el('framelabel').textContent =
|
||||||
`f ${f} / ${state.dense.length - 1} · ${(f / state.fps).toFixed(2)}s · plate f${pf}` +
|
`f ${f} / ${state.dense.length - 1} · ${(f / state.fps).toFixed(2)}s · ` +
|
||||||
|
`plate f${pf} · mouth f${leadIndex(f)}` +
|
||||||
|
(lead ? ` (${lead > 0 ? '+' : ''}${lead} = ${(lead / state.fps * 1000).toFixed(0)}ms)` : '') +
|
||||||
(state.keep.has(f) ? ' · KEPT' : ' · held');
|
(state.keep.has(f) ? ' · KEPT' : ' · held');
|
||||||
|
|
||||||
const c1 = el('cv-source'), g1 = c1.getContext('2d');
|
const c1 = el('cv-source'), g1 = c1.getContext('2d');
|
||||||
|
|
@ -301,11 +419,43 @@ function drawPanes() {
|
||||||
g2.moveTo(c2.width / 2, 0); g2.lineTo(c2.width / 2, c2.height);
|
g2.moveTo(c2.width / 2, 0); g2.lineTo(c2.width / 2, c2.height);
|
||||||
g2.moveTo(0, c2.height / 2); g2.lineTo(c2.width, c2.height / 2); g2.stroke();
|
g2.moveTo(0, c2.height / 2); g2.lineTo(c2.width, c2.height / 2); g2.stroke();
|
||||||
const z = (pts) => pts.map((p) => ({ x: p.x * ZOOM, y: p.y * ZOOM }));
|
const z = (pts) => pts.map((p) => ({ x: p.x * ZOOM, y: p.y * ZOOM }));
|
||||||
|
const mf = leadIndex(f);
|
||||||
strokePts(g2, z(state.plates[pf]), '#3b4a63');
|
strokePts(g2, z(state.plates[pf]), '#3b4a63');
|
||||||
strokePts(g2, z(state.outer[f]), '#4ade80');
|
// With a lead set, the unshifted contour is drawn as a ghost so the offset is
|
||||||
if (!state.hidden[f]) strokePts(g2, z(state.inner[f]), '#f87171');
|
// something you can see rather than something you have to trust.
|
||||||
|
if (mf !== f) strokePts(g2, z(state.outer[f]), '#2f6b46');
|
||||||
|
strokePts(g2, z(state.outer[mf]), '#4ade80');
|
||||||
|
if (!state.hidden[mf]) strokePts(g2, z(state.inner[mf]), '#f87171');
|
||||||
|
|
||||||
compositeRender(el('cv-render'), f, ZOOM);
|
compositeRender(el('cv-render'), f, ZOOM);
|
||||||
|
drawInteriorDebug(f);
|
||||||
|
}
|
||||||
|
|
||||||
|
// What the teeth measurement actually saw: sampled region, pixels above
|
||||||
|
// threshold in green, the resolved line in amber. Recomputed for the current
|
||||||
|
// frame only, so it costs nothing to keep on screen.
|
||||||
|
function drawInteriorDebug(fRaw) {
|
||||||
|
const f = leadIndex(fRaw);
|
||||||
|
const host = el('cv-teeth');
|
||||||
|
const img = state.images[f];
|
||||||
|
if (!img || state.hidden[f]) {
|
||||||
|
host.innerHTML = '';
|
||||||
|
el('teethinfo').textContent = state.images.length ? 'mouth closed' : 'no source frames';
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
const o = opts();
|
||||||
|
const ctx = document.createElement('canvas').getContext('2d', { willReadFrequently: true });
|
||||||
|
const m = extractTeeth(img, LIPS_INNER.map((k) => state.dense[f][k]), ctx, o, true);
|
||||||
|
host.innerHTML = '';
|
||||||
|
if (m.debug) {
|
||||||
|
m.debug.style.width = '170px';
|
||||||
|
m.debug.style.imageRendering = 'pixelated';
|
||||||
|
host.append(m.debug);
|
||||||
|
}
|
||||||
|
const te = state.teeth[f];
|
||||||
|
el('teethinfo').textContent =
|
||||||
|
`contrast ${m.contrast.toFixed(3)} / gate ${o.teethOn.toFixed(2)} · ` +
|
||||||
|
`area ${m.area}px · ${te.show ? 'SHOWN' : 'hidden'}`;
|
||||||
}
|
}
|
||||||
|
|
||||||
function strokePts(g, pts, color, lw = 1) {
|
function strokePts(g, pts, color, lw = 1) {
|
||||||
|
|
@ -400,11 +550,21 @@ function exportTake() {
|
||||||
// Sparse: one plate key per frame a human draws.
|
// Sparse: one plate key per frame a human draws.
|
||||||
{ name: 'head', kind: 'plate', z: 0, interp: 'hold',
|
{ name: 'head', kind: 'plate', z: 0, interp: 'hold',
|
||||||
keys: kept.map((f, i) => ({ f: i, plate: i, src: f })) },
|
keys: kept.map((f, i) => ({ f: i, plate: i, src: f })) },
|
||||||
// Dense: the traced mouth keeps every frame.
|
// Dense: the traced mouth keeps every frame. The lead is baked in here -
|
||||||
|
// key f carries the pose from source frame f+lead - so the renderer never
|
||||||
|
// needs to know about it.
|
||||||
{ name: 'mouth', kind: 'poly', z: 30, color: 'skin_dark', interp: 'hold',
|
{ name: 'mouth', kind: 'poly', z: 30, color: 'skin_dark', interp: 'hold',
|
||||||
keys: state.outer.map((pts, f) => ({ f, src: f, pts })) },
|
keys: state.outer.map((_, f) => ({ f, src: leadIndex(f), pts: state.outer[leadIndex(f)] })) },
|
||||||
{ name: 'mouth_in', kind: 'poly', z: 31, color: 'mouth_dark', interp: 'hold', parent: 'mouth',
|
{ name: 'mouth_in', kind: 'poly', z: 31, color: 'mouth_dark', interp: 'hold', parent: 'mouth',
|
||||||
keys: state.inner.map((pts, f) => (state.hidden[f] ? { f, hidden: true } : { f, src: f, pts })) },
|
keys: state.inner.map((_, f) => {
|
||||||
|
const m = leadIndex(f);
|
||||||
|
return state.hidden[m] ? { f, hidden: true } : { f, src: m, pts: state.inner[m] };
|
||||||
|
}) },
|
||||||
|
{ name: 'teeth', kind: 'poly', z: 32, color: 'teeth', interp: 'hold', parent: 'mouth_in',
|
||||||
|
keys: state.teeth.map((_, f) => {
|
||||||
|
const m = leadIndex(f), te = state.teeth[m];
|
||||||
|
return te.show && te.pts ? { f, src: m, pts: te.pts } : { f, hidden: true };
|
||||||
|
}) },
|
||||||
],
|
],
|
||||||
};
|
};
|
||||||
const text = writeTake(take)
|
const text = writeTake(take)
|
||||||
|
|
@ -440,11 +600,15 @@ async function runFrames() {
|
||||||
}
|
}
|
||||||
const { dense, missing } = await detectAll(images);
|
const { dense, missing } = await detectAll(images);
|
||||||
state.images = images; state.dense = dense;
|
state.images = images; state.dense = dense;
|
||||||
|
state.aspect = images[0].naturalWidth / images[0].naturalHeight;
|
||||||
|
status('measuring mouth interiors…');
|
||||||
|
state.interior = measureAll(images, dense, opts());
|
||||||
el('scrub').max = dense.length - 1;
|
el('scrub').max = dense.length - 1;
|
||||||
state.frame = 0;
|
state.frame = 0;
|
||||||
rebuild(true);
|
rebuild(true);
|
||||||
const dur = (dense.length / state.fps).toFixed(2);
|
const dur = (dense.length / state.fps).toFixed(2);
|
||||||
status(`${images.length} frames · ${state.fps}fps · ${dur}s` +
|
status(`${images.length} frames · ${images[0].naturalWidth}x${images[0].naturalHeight} · ` +
|
||||||
|
`${state.fps}fps · ${dur}s` +
|
||||||
(state.audio ? ' · audio loaded' : ' · no audio') +
|
(state.audio ? ' · audio loaded' : ' · no audio') +
|
||||||
(missing.length ? ` · no face on ${missing.length} (held previous)` : ''),
|
(missing.length ? ` · no face on ${missing.length} (held previous)` : ''),
|
||||||
missing.length ? 'warn' : 'ok');
|
missing.length ? 'warn' : 'ok');
|
||||||
|
|
@ -455,6 +619,10 @@ function runSynthetic() {
|
||||||
state.images = [];
|
state.images = [];
|
||||||
attachAudio(null);
|
attachAudio(null);
|
||||||
state.fps = 12;
|
state.fps = 12;
|
||||||
|
state.aspect = 1; // synthetic landmarks are generated square
|
||||||
|
state.lead = 0;
|
||||||
|
state.interior = null; // no pixels, so no teeth
|
||||||
|
|
||||||
state.dense = synthDense(72);
|
state.dense = synthDense(72);
|
||||||
el('scrub').max = 71;
|
el('scrub').max = 71;
|
||||||
state.frame = 0;
|
state.frame = 0;
|
||||||
|
|
@ -462,10 +630,16 @@ function runSynthetic() {
|
||||||
status('synthetic — exercises everything below detection', 'ok');
|
status('synthetic — exercises everything below detection', 'ok');
|
||||||
}
|
}
|
||||||
|
|
||||||
for (const id of ['verts', 'smoothWin', 'contourSmooth', 'apertureThresh', 'tol']) {
|
for (const id of ['verts', 'smoothWin', 'contourSmooth', 'apertureThresh', 'tol',
|
||||||
|
'teethOn', 'teethDwell', 'teethErode', 'tongueReject', 'blobGrow',
|
||||||
|
'topBias', 'teethVerts', 'teethSmooth', 'lead']) {
|
||||||
el(id).addEventListener('input', () => {
|
el(id).addEventListener('input', () => {
|
||||||
el(id + 'v').textContent = id === 'apertureThresh' || id === 'tol'
|
el(id + 'v').textContent = id === 'apertureThresh' || id === 'tol'
|
||||||
? (+el(id).value / 1000).toFixed(3) : el(id).value;
|
? (+el(id).value / 1000).toFixed(3)
|
||||||
|
: ['teethOn', 'teethErode', 'tongueReject', 'topBias'].includes(id)
|
||||||
|
? (+el(id).value / 100).toFixed(2)
|
||||||
|
: id === 'lead' && +el(id).value > 0 ? `+${el(id).value}`
|
||||||
|
: el(id).value;
|
||||||
if (id === 'tol') return; // tol only matters when you ask for a suggestion
|
if (id === 'tol') return; // tol only matters when you ask for a suggestion
|
||||||
rebuild(false);
|
rebuild(false);
|
||||||
});
|
});
|
||||||
|
|
@ -533,6 +707,12 @@ window.addEventListener('keydown', (e) => {
|
||||||
else if (e.key === 'ArrowLeft') { seekTo(Math.max(0, state.frame - 1)); drawAll(); }
|
else if (e.key === 'ArrowLeft') { seekTo(Math.max(0, state.frame - 1)); drawAll(); }
|
||||||
else if (e.key === 'Backspace' || e.key === 'Delete' || e.key === 'x') {
|
else if (e.key === 'Backspace' || e.key === 'Delete' || e.key === 'x') {
|
||||||
state.keep.delete(state.frame === 0 ? -1 : state.frame); drawAll();
|
state.keep.delete(state.frame === 0 ? -1 : state.frame); drawAll();
|
||||||
|
} else if (e.key === '[' || e.key === ']') {
|
||||||
|
const n = el('lead');
|
||||||
|
n.value = Math.max(+n.min, Math.min(+n.max, +n.value + (e.key === ']' ? 1 : -1)));
|
||||||
|
state.lead = +n.value;
|
||||||
|
el('leadv').textContent = state.lead > 0 ? `+${state.lead}` : String(state.lead);
|
||||||
|
drawAll();
|
||||||
} else if (e.key === 'b') {
|
} else if (e.key === 'b') {
|
||||||
const sel = el('plateMode');
|
const sel = el('plateMode');
|
||||||
sel.selectedIndex = (sel.selectedIndex + 1) % sel.options.length;
|
sel.selectedIndex = (sel.selectedIndex + 1) % sel.options.length;
|
||||||
|
|
@ -585,6 +765,11 @@ PALETTE.forEach((p) => {
|
||||||
|
|
||||||
// #synth / #frames autorun, so the tool can be driven headlessly for smoke tests
|
// #synth / #frames autorun, so the tool can be driven headlessly for smoke tests
|
||||||
// and deep-linked. Detection needs WebGL; the synthetic path does not.
|
// and deep-linked. Detection needs WebGL; the synthetic path does not.
|
||||||
|
window.addEventListener('error', (e) => {
|
||||||
|
const s = document.getElementById('status');
|
||||||
|
if (s) { s.textContent = e.message; s.className = 'err'; }
|
||||||
|
});
|
||||||
|
|
||||||
if (location.hash === '#synth') runSynthetic();
|
if (location.hash === '#synth') runSynthetic();
|
||||||
else if (location.hash === '#frames') runFrames();
|
else if (location.hash === '#frames') runFrames();
|
||||||
else status('ready — Load frames, then step with \u2190 \u2192 and delete with X');
|
else status('ready — Load frames, then step with \u2190 \u2192 and delete with X');
|
||||||
|
|
|
||||||
225
js/interior.js
Normal file
225
js/interior.js
Normal file
|
|
@ -0,0 +1,225 @@
|
||||||
|
// Mouth interior from image content.
|
||||||
|
//
|
||||||
|
// MediaPipe has no landmarks inside the lips - the inner ring bounds the cavity
|
||||||
|
// and everything within it is just pixels. So teeth come from the picture.
|
||||||
|
//
|
||||||
|
// The hazard is vertex correspondence. A traced contour reorders between frames
|
||||||
|
// and boils, which is the failure docs/design.md exists to avoid. The way
|
||||||
|
// out for a blob specifically is RADIAL SAMPLING: march outward from the
|
||||||
|
// centroid along N fixed directions and take the last pixel inside. Vertex k is
|
||||||
|
// then always "the blob's extent in direction k" - correspondence holds by
|
||||||
|
// construction, the vertex count is fixed, and the result smooths over time
|
||||||
|
// without any reordering being possible. It also yields a star-shaped
|
||||||
|
// reduction, which is what flat blocks of colour want anyway.
|
||||||
|
|
||||||
|
export function otsuForTest(h, t) { return otsu(h, t); }
|
||||||
|
|
||||||
|
// Otsu's threshold plus its two class means. The means matter as much as the
|
||||||
|
// threshold: Otsu ALWAYS returns a split, including on a homogeneous region, so
|
||||||
|
// their separation is the only thing that says the split means anything.
|
||||||
|
function otsu(hist, total) {
|
||||||
|
let sum = 0;
|
||||||
|
for (let i = 0; i < 256; i++) sum += i * hist[i];
|
||||||
|
let sumB = 0, wB = 0, best = 0, bestVar = -1, bestDark = 0, bestBright = 0;
|
||||||
|
for (let t = 0; t < 256; t++) {
|
||||||
|
wB += hist[t];
|
||||||
|
if (!wB) continue;
|
||||||
|
const wF = total - wB;
|
||||||
|
if (!wF) break;
|
||||||
|
sumB += t * hist[t];
|
||||||
|
const mDark = sumB / wB, mBright = (sum - sumB) / wF;
|
||||||
|
const between = wB * wF * (mDark - mBright) * (mDark - mBright);
|
||||||
|
if (between > bestVar) { bestVar = between; best = t; bestDark = mDark; bestBright = mBright; }
|
||||||
|
}
|
||||||
|
return { thr: best, mDark: bestDark, mBright: bestBright };
|
||||||
|
}
|
||||||
|
|
||||||
|
const pointInPoly = (pts, x, y) => {
|
||||||
|
let inside = false;
|
||||||
|
for (let i = 0, j = pts.length - 1; i < pts.length; j = i++) {
|
||||||
|
if ((pts[i].y > y) !== (pts[j].y > y) &&
|
||||||
|
x < ((pts[j].x - pts[i].x) * (y - pts[i].y)) / (pts[j].y - pts[i].y) + pts[i].x) inside = !inside;
|
||||||
|
}
|
||||||
|
return inside;
|
||||||
|
};
|
||||||
|
|
||||||
|
// Shrink or grow a ring about its centroid. MediaPipe's inner lip landmarks sit
|
||||||
|
// slightly OUTSIDE the real opening, so sampling the ring as given includes lip
|
||||||
|
// pixels - bright, and right at the boundary where they do most damage.
|
||||||
|
export function scaleRing(pts, k) {
|
||||||
|
let cx = 0, cy = 0;
|
||||||
|
for (const p of pts) { cx += p.x; cy += p.y; }
|
||||||
|
cx /= pts.length; cy /= pts.length;
|
||||||
|
return pts.map((p) => ({ x: cx + (p.x - cx) * k, y: cy + (p.y - cy) * k }));
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ---- binary morphology on the candidate mask ---- */
|
||||||
|
|
||||||
|
function erodeMask(m, w, h) {
|
||||||
|
const o = new Uint8Array(m.length);
|
||||||
|
for (let y = 1; y < h - 1; y++) for (let x = 1; x < w - 1; x++) {
|
||||||
|
const i = y * w + x;
|
||||||
|
o[i] = m[i] && m[i - 1] && m[i + 1] && m[i - w] && m[i + w] ? 1 : 0;
|
||||||
|
}
|
||||||
|
return o;
|
||||||
|
}
|
||||||
|
function dilateMask(m, w, h) {
|
||||||
|
const o = new Uint8Array(m.length);
|
||||||
|
for (let y = 1; y < h - 1; y++) for (let x = 1; x < w - 1; x++) {
|
||||||
|
const i = y * w + x;
|
||||||
|
o[i] = m[i] || m[i - 1] || m[i + 1] || m[i - w] || m[i + w] ? 1 : 0;
|
||||||
|
}
|
||||||
|
return o;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Largest 4-connected component, scored with a bias toward the TOP of the
|
||||||
|
// cavity: upper teeth hang from the lip, and the usual false positive is the
|
||||||
|
// tongue sitting lower down. Area alone picks the tongue when the mouth is wide.
|
||||||
|
function bestComponent(mask, w, h, topBias) {
|
||||||
|
const label = new Int32Array(mask.length).fill(-1);
|
||||||
|
const stack = [];
|
||||||
|
let best = null, id = 0;
|
||||||
|
for (let s = 0; s < mask.length; s++) {
|
||||||
|
if (!mask[s] || label[s] >= 0) continue;
|
||||||
|
stack.length = 0; stack.push(s);
|
||||||
|
label[s] = id;
|
||||||
|
const px = [];
|
||||||
|
let sumY = 0;
|
||||||
|
while (stack.length) {
|
||||||
|
const i = stack.pop();
|
||||||
|
px.push(i);
|
||||||
|
sumY += (i / w) | 0;
|
||||||
|
const x = i % w, y = (i / w) | 0;
|
||||||
|
if (x > 0 && mask[i - 1] && label[i - 1] < 0) { label[i - 1] = id; stack.push(i - 1); }
|
||||||
|
if (x < w - 1 && mask[i + 1] && label[i + 1] < 0) { label[i + 1] = id; stack.push(i + 1); }
|
||||||
|
if (y > 0 && mask[i - w] && label[i - w] < 0) { label[i - w] = id; stack.push(i - w); }
|
||||||
|
if (y < h - 1 && mask[i + w] && label[i + w] < 0) { label[i + w] = id; stack.push(i + w); }
|
||||||
|
}
|
||||||
|
const meanY = sumY / px.length / h; // 0 top, 1 bottom
|
||||||
|
const score = px.length * (1 - topBias * meanY);
|
||||||
|
if (!best || score > best.score) best = { score, px, area: px.length, meanY };
|
||||||
|
id++;
|
||||||
|
}
|
||||||
|
return best;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Radial sampling from the centroid: N fixed directions, last pixel inside.
|
||||||
|
function radialContour(mask, w, h, cx, cy, n) {
|
||||||
|
const pts = [];
|
||||||
|
const maxR = Math.hypot(w, h);
|
||||||
|
let prev = 1;
|
||||||
|
for (let k = 0; k < n; k++) {
|
||||||
|
const a = -(k / n) * Math.PI * 2; // slot 0 = +x, 5/20 = top
|
||||||
|
const dx = Math.cos(a), dy = Math.sin(a);
|
||||||
|
let hit = 0;
|
||||||
|
for (let r = 0.5; r < maxR; r += 0.5) {
|
||||||
|
const x = Math.round(cx + dx * r), y = Math.round(cy + dy * r);
|
||||||
|
if (x < 0 || y < 0 || x >= w || y >= h) break;
|
||||||
|
if (mask[y * w + x]) hit = r;
|
||||||
|
else if (hit > 0 && r > hit + 2) break; // tolerate a 2px gap, then stop
|
||||||
|
}
|
||||||
|
// A ray that escapes immediately would collapse the polygon; hold the last
|
||||||
|
// good radius so the shape stays closed rather than spiking to the centre.
|
||||||
|
if (hit <= 0) hit = prev * 0.6;
|
||||||
|
prev = hit;
|
||||||
|
pts.push({ x: cx + dx * hit, y: cy + dy * hit });
|
||||||
|
}
|
||||||
|
return pts;
|
||||||
|
}
|
||||||
|
|
||||||
|
/* ---- the extraction ---- */
|
||||||
|
|
||||||
|
export function extractTeeth(img, innerNorm, ctx, o, wantDebug = false) {
|
||||||
|
const none = { contour: null, contrast: 0, area: 0, debug: null };
|
||||||
|
const ring = scaleRing(innerNorm, 1 - (o.cavityErode ?? 0.18));
|
||||||
|
|
||||||
|
let x0 = 1, y0 = 1, x1 = 0, y1 = 0;
|
||||||
|
for (const p of ring) {
|
||||||
|
x0 = Math.min(x0, p.x); y0 = Math.min(y0, p.y);
|
||||||
|
x1 = Math.max(x1, p.x); y1 = Math.max(y1, p.y);
|
||||||
|
}
|
||||||
|
const W = img.naturalWidth, H = img.naturalHeight;
|
||||||
|
const px0 = Math.max(0, Math.floor(x0 * W)), py0 = Math.max(0, Math.floor(y0 * H));
|
||||||
|
const pw = Math.min(W - px0, Math.ceil((x1 - x0) * W)), ph = Math.min(H - py0, Math.ceil((y1 - y0) * H));
|
||||||
|
if (pw < 5 || ph < 5) return none;
|
||||||
|
|
||||||
|
ctx.canvas.width = pw; ctx.canvas.height = ph;
|
||||||
|
ctx.drawImage(img, px0, py0, pw, ph, 0, 0, pw, ph);
|
||||||
|
const src = ctx.getImageData(0, 0, pw, ph);
|
||||||
|
const d = src.data;
|
||||||
|
|
||||||
|
const poly = ring.map((p) => ({ x: p.x * W - px0, y: p.y * H - py0 }));
|
||||||
|
const hist = new Uint32Array(256);
|
||||||
|
const lum = new Float32Array(pw * ph);
|
||||||
|
const red = new Float32Array(pw * ph);
|
||||||
|
const inReg = new Uint8Array(pw * ph);
|
||||||
|
let n = 0;
|
||||||
|
for (let y = 0; y < ph; y++) for (let x = 0; x < pw; x++) {
|
||||||
|
if (!pointInPoly(poly, x + 0.5, y + 0.5)) continue;
|
||||||
|
const i = y * pw + x, oo = i * 4;
|
||||||
|
const R = d[oo], G = d[oo + 1], B = d[oo + 2];
|
||||||
|
lum[i] = (0.299 * R + 0.587 * G + 0.114 * B) | 0;
|
||||||
|
// Tongue is red relative to its own brightness; teeth are near-neutral.
|
||||||
|
red[i] = (R - (G + B) / 2) / 255;
|
||||||
|
inReg[i] = 1; hist[lum[i]]++; n++;
|
||||||
|
}
|
||||||
|
if (n < 24) return none;
|
||||||
|
|
||||||
|
const { thr, mDark, mBright } = otsu(hist, n);
|
||||||
|
const contrast = (mBright - mDark) / 255;
|
||||||
|
|
||||||
|
let mask = new Uint8Array(pw * ph);
|
||||||
|
for (let i = 0; i < mask.length; i++) {
|
||||||
|
mask[i] = inReg[i] && lum[i] > thr && red[i] < (o.tongueReject ?? 0.18) ? 1 : 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
// Open once to despeckle, then apply the signed size adjustment.
|
||||||
|
mask = dilateMask(erodeMask(mask, pw, ph), pw, ph);
|
||||||
|
const grow = o.blobGrow | 0;
|
||||||
|
for (let k = 0; k < Math.abs(grow); k++) {
|
||||||
|
mask = grow < 0 ? erodeMask(mask, pw, ph) : dilateMask(mask, pw, ph);
|
||||||
|
}
|
||||||
|
|
||||||
|
const comp = bestComponent(mask, pw, ph, o.topBias ?? 0.6);
|
||||||
|
if (!comp || comp.area < (o.minArea ?? 12)) {
|
||||||
|
return { contour: null, contrast, area: comp ? comp.area : 0,
|
||||||
|
debug: wantDebug ? debugCanvas(src, inReg, mask, pw, ph, null) : null };
|
||||||
|
}
|
||||||
|
|
||||||
|
const only = new Uint8Array(mask.length);
|
||||||
|
let cx = 0, cy = 0;
|
||||||
|
for (const i of comp.px) { only[i] = 1; cx += i % pw; cy += (i / pw) | 0; }
|
||||||
|
cx /= comp.px.length; cy /= comp.px.length;
|
||||||
|
|
||||||
|
const local = radialContour(only, pw, ph, cx, cy, o.teethVerts ?? 10);
|
||||||
|
const contour = local.map((p) => ({ x: (p.x + px0) / W, y: (p.y + py0) / H }));
|
||||||
|
|
||||||
|
return {
|
||||||
|
contour, contrast, area: comp.area,
|
||||||
|
debug: wantDebug ? debugCanvas(src, inReg, only, pw, ph, local) : null,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
// Sampled region dimmed, kept pixels green, extracted contour in amber.
|
||||||
|
function debugCanvas(src, inReg, mask, pw, ph, local) {
|
||||||
|
const c = document.createElement('canvas');
|
||||||
|
c.width = pw; c.height = ph;
|
||||||
|
const g = c.getContext('2d');
|
||||||
|
const out = new ImageData(pw, ph);
|
||||||
|
for (let i = 0; i < pw * ph; i++) {
|
||||||
|
const o = i * 4;
|
||||||
|
const [r, gr, b] = [src.data[o], src.data[o + 1], src.data[o + 2]];
|
||||||
|
if (!inReg[i]) { out.data[o] = r * 0.25; out.data[o + 1] = gr * 0.25; out.data[o + 2] = b * 0.25; }
|
||||||
|
else if (mask[i]) { out.data[o] = 60; out.data[o + 1] = 230; out.data[o + 2] = 120; }
|
||||||
|
else { out.data[o] = r; out.data[o + 1] = gr; out.data[o + 2] = b; }
|
||||||
|
out.data[o + 3] = 255;
|
||||||
|
}
|
||||||
|
g.putImageData(out, 0, 0);
|
||||||
|
if (local && local.length) {
|
||||||
|
g.strokeStyle = '#fbbf24'; g.lineWidth = 1;
|
||||||
|
g.beginPath();
|
||||||
|
local.forEach((p, i) => (i ? g.lineTo(p.x, p.y) : g.moveTo(p.x, p.y)));
|
||||||
|
g.closePath(); g.stroke();
|
||||||
|
}
|
||||||
|
return c;
|
||||||
|
}
|
||||||
|
|
@ -1,7 +1,7 @@
|
||||||
// MediaPipe FaceLandmarker index tables.
|
// MediaPipe FaceLandmarker index tables.
|
||||||
// Ring arrays are ORDERED traversals, not raw connection sets: vertex position
|
// Ring arrays are ORDERED traversals, not raw connection sets: vertex position
|
||||||
// within a ring is the vertex's identity, and every downstream stage depends on
|
// within a ring is the vertex's identity, and every downstream stage depends on
|
||||||
// that ordering being stable. See docs/roto-puppet.md, "Fixed topology".
|
// that ordering being stable. See docs/design.md, "Fixed topology".
|
||||||
|
|
||||||
// Rigid landmarks for the similarity fit. Eye corners, nose bridge, nose tip.
|
// Rigid landmarks for the similarity fit. Eye corners, nose bridge, nose tip.
|
||||||
// Nothing here may be a feature that moves under performance: including the
|
// Nothing here may be a feature that moves under performance: including the
|
||||||
|
|
|
||||||
|
|
@ -1,17 +1,26 @@
|
||||||
// Analysis: dense track -> stabilised head-local contours -> selected keys.
|
// Analysis: dense track -> stabilised head-local contours -> selected keys.
|
||||||
// All policy lives here, never in the renderer. See docs/roto-puppet.md,
|
// All policy lives here, never in the renderer. See docs/design.md,
|
||||||
// "The take is the contract".
|
// "The take is the contract".
|
||||||
|
|
||||||
import { RIGID, LIPS_OUTER, LIPS_INNER, APERTURE, FACE_OVAL, EYE_INNER, subsampleSlots } from './landmarks.js';
|
import { RIGID, LIPS_OUTER, LIPS_INNER, APERTURE, FACE_OVAL, EYE_INNER, subsampleSlots } from './landmarks.js';
|
||||||
import { fitSimilarity, applySimAll, applySim, fitResidual, procrustesMean, smoothTransforms, movingAverage } from './mathutil.js';
|
import { fitSimilarity, applySimAll, applySim, fitResidual, procrustesMean, smoothTransforms, movingAverage } from './mathutil.js';
|
||||||
|
|
||||||
const pick = (lm, idx) => idx.map((i) => ({ x: lm[i].x, y: lm[i].y }));
|
// MediaPipe normalises x by image WIDTH and y by image HEIGHT, so its normalised
|
||||||
|
// space is anisotropic: for a 1080x1920 frame, one unit of x is 1080px and one
|
||||||
|
// unit of y is 1920px. Treating those as comparable stretches everything
|
||||||
|
// horizontally by H/W, and worse, makes fitSimilarity fit a "rotation" in a
|
||||||
|
// sheared space, so head roll comes out subtly wrong as well.
|
||||||
|
//
|
||||||
|
// Multiplying x by aspect = W/H converts to an ISOTROPIC space whose unit is one
|
||||||
|
// image height, so equal numbers mean equal pixels. Everything downstream -
|
||||||
|
// Procrustes, the similarity fit, the raster transform - depends on that.
|
||||||
|
const pick = (lm, idx, aspect) => idx.map((i) => ({ x: lm[i].x * aspect, y: lm[i].y }));
|
||||||
|
|
||||||
// Stage 1-3: fit the rigid transform per frame, smooth its parameters, then map
|
// Stage 1-3: fit the rigid transform per frame, smooth its parameters, then map
|
||||||
// every contour through it into the reference frame. The result is head-local:
|
// every contour through it into the reference frame. The result is head-local:
|
||||||
// translation, roll and depth-scale of the head are gone.
|
// translation, roll and depth-scale of the head are gone.
|
||||||
export function stabilize(dense, smoothRadius) {
|
export function stabilize(dense, smoothRadius, aspect = 1) {
|
||||||
const rigid = dense.map((f) => pick(f, RIGID));
|
const rigid = dense.map((f) => pick(f, RIGID, aspect));
|
||||||
const ref = procrustesMean(rigid);
|
const ref = procrustesMean(rigid);
|
||||||
const raw = rigid.map((r) => fitSimilarity(r, ref));
|
const raw = rigid.map((r) => fitSimilarity(r, ref));
|
||||||
const tfs = smoothTransforms(raw, smoothRadius);
|
const tfs = smoothTransforms(raw, smoothRadius);
|
||||||
|
|
@ -26,12 +35,12 @@ export function stabilize(dense, smoothRadius) {
|
||||||
// Residual rises with out-of-plane rotation, which no 2D similarity can
|
// Residual rises with out-of-plane rotation, which no 2D similarity can
|
||||||
// remove. High values mean this section wants a different head plate.
|
// remove. High values mean this section wants a different head plate.
|
||||||
residual: tfs.map((tf, i) => fitResidual(tf, rigid[i], ref)),
|
residual: tfs.map((tf, i) => fitResidual(tf, rigid[i], ref)),
|
||||||
outer: dense.map((f, i) => applySimAll(tfs[i], pick(f, LIPS_OUTER))),
|
outer: dense.map((f, i) => applySimAll(tfs[i], pick(f, LIPS_OUTER, aspect))),
|
||||||
inner: dense.map((f, i) => applySimAll(tfs[i], pick(f, LIPS_INNER))),
|
inner: dense.map((f, i) => applySimAll(tfs[i], pick(f, LIPS_INNER, aspect))),
|
||||||
oval: dense.map((f, i) => applySimAll(tfs[i], pick(f, FACE_OVAL))),
|
oval: dense.map((f, i) => applySimAll(tfs[i], pick(f, FACE_OVAL, aspect))),
|
||||||
eyes: dense.map((f, i) => applySimAll(tfs[i], pick(f, EYE_INNER))),
|
eyes: dense.map((f, i) => applySimAll(tfs[i], pick(f, EYE_INNER, aspect))),
|
||||||
aperture: dense.map((f, i) => {
|
aperture: dense.map((f, i) => {
|
||||||
const a = applySimAll(tfs[i], pick(f, APERTURE));
|
const a = applySimAll(tfs[i], pick(f, APERTURE, aspect));
|
||||||
return Math.hypot(a[0].x - a[1].x, a[0].y - a[1].y);
|
return Math.hypot(a[0].x - a[1].x, a[0].y - a[1].y);
|
||||||
}),
|
}),
|
||||||
};
|
};
|
||||||
|
|
@ -113,7 +122,7 @@ export function activeKey(keys, f) {
|
||||||
|
|
||||||
// Temporal smoothing of a contour, per vertex, across time.
|
// Temporal smoothing of a contour, per vertex, across time.
|
||||||
//
|
//
|
||||||
// docs/roto-puppet.md says to smooth the transform and never the contour. That
|
// docs/design.md says to smooth the transform and never the contour. That
|
||||||
// was correct while keys were sparse: sampling at velocity minima rejected
|
// was correct while keys were sparse: sampling at velocity minima rejected
|
||||||
// per-frame detector noise for free. With a key on every frame the noise is
|
// per-frame detector noise for free. With a key on every frame the noise is
|
||||||
// visible as a shimmer along the lip edge, so a bounded exception applies -
|
// visible as a shimmer along the lip edge, so a bounded exception applies -
|
||||||
|
|
@ -168,3 +177,12 @@ export function heldFrame(kept, f) {
|
||||||
for (const k of kept) { if (k <= f) hit = k; else break; }
|
for (const k of kept) { if (k <= f) hit = k; else break; }
|
||||||
return hit;
|
return hit;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Shift a performance track against the clock, clamped at the ends.
|
||||||
|
//
|
||||||
|
// Pure and exported so the shift can actually be asserted: "the slider feels
|
||||||
|
// like it does nothing" is otherwise indistinguishable from "the slider does
|
||||||
|
// nothing", and at 24fps a lead of 1 is 42ms, which is small enough to doubt.
|
||||||
|
export function shiftIndex(f, lead, n) {
|
||||||
|
return Math.min(n - 1, Math.max(0, f + lead));
|
||||||
|
}
|
||||||
|
|
|
||||||
145
js/selftest.js
145
js/selftest.js
|
|
@ -2,16 +2,17 @@
|
||||||
// module graph the tool uses is what gets tested.
|
// module graph the tool uses is what gets tested.
|
||||||
//
|
//
|
||||||
// The ring-simplicity check exists because "fixed topology" is load-bearing in
|
// The ring-simplicity check exists because "fixed topology" is load-bearing in
|
||||||
// docs/roto-puppet.md: because hold parts CUT between poses rather than
|
// docs/design.md: because hold parts CUT between poses rather than
|
||||||
// interpolating, a ring whose vertex order is wrong self-intersects and renders
|
// interpolating, a ring whose vertex order is wrong self-intersects and renders
|
||||||
// as blocks meeting at corners. It is invisible at some vertex counts and obvious
|
// as blocks meeting at corners. It is invisible at some vertex counts and obvious
|
||||||
// at others, so it needs an assertion rather than an eyeball.
|
// at others, so it needs an assertion rather than an eyeball.
|
||||||
|
|
||||||
import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, subsampleSlots, subsampleRing } from './landmarks.js';
|
import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, subsampleSlots, subsampleRing } from './landmarks.js';
|
||||||
import { fitSimilarity, applySim, procrustesMean, smoothTransforms } from './mathutil.js';
|
import { fitSimilarity, applySim, procrustesMean, smoothTransforms } from './mathutil.js';
|
||||||
import { stabilize, toRasterRing, selectKeys, activeKey } from './pipeline.js';
|
import { stabilize, toRasterRing, selectKeys, activeKey, shiftIndex } from './pipeline.js';
|
||||||
import { IndexedRaster, hexToRgb } from './raster.js';
|
import { IndexedRaster, hexToRgb } from './raster.js';
|
||||||
import { writeTake } from './take.js';
|
import { writeTake } from './take.js';
|
||||||
|
import { otsuForTest, scaleRing } from './interior.js';
|
||||||
import { synthDense } from './synth.js';
|
import { synthDense } from './synth.js';
|
||||||
|
|
||||||
const results = [];
|
const results = [];
|
||||||
|
|
@ -44,6 +45,50 @@ const spreadX = (frames, slot) => {
|
||||||
|
|
||||||
/* ---- the tests ---- */
|
/* ---- the tests ---- */
|
||||||
|
|
||||||
|
// Cross-check every el('id') in app.js against the ids in index.html.
|
||||||
|
//
|
||||||
|
// This bug class has bitten twice: a knob wired in app.js but absent from the
|
||||||
|
// markup throws during wiring, which aborts the rest of the module and leaves a
|
||||||
|
// blank page. The symptom ("nothing happens") points nowhere near the cause, so
|
||||||
|
// it is worth an automated check rather than vigilance.
|
||||||
|
export async function runWiring() {
|
||||||
|
const out = [];
|
||||||
|
try {
|
||||||
|
const [app, html] = await Promise.all([
|
||||||
|
fetch('./js/app.js').then((r) => r.text()),
|
||||||
|
fetch('./index.html').then((r) => r.text()),
|
||||||
|
]);
|
||||||
|
const ids = new Set([...app.matchAll(/\bel\(\s*['"]([\w-]+)['"]\s*\)/g)].map((m) => m[1]));
|
||||||
|
const list = app.match(/for \(const id of \[([\s\S]*?)\]\)/);
|
||||||
|
if (list) {
|
||||||
|
for (const m of list[1].matchAll(/'([\w]+)'/g)) { ids.add(m[1]); ids.add(m[1] + 'v'); }
|
||||||
|
}
|
||||||
|
// Duplicate keys in an object literal are silent in JS - the last one wins.
|
||||||
|
// In opts() that meant the teeth vertex slider was quietly driving the lip
|
||||||
|
// vertex count while the lip slider did nothing at all.
|
||||||
|
const lit = app.match(/const opts = \(\) => \(\{([\s\S]*?)\n\}\);/);
|
||||||
|
if (lit) {
|
||||||
|
const keys = [...lit[1].matchAll(/^\s*([A-Za-z_$][\w$]*)\s*:/gm)].map((m) => m[1]);
|
||||||
|
const dupes = keys.filter((k, i) => keys.indexOf(k) !== i);
|
||||||
|
out.push({ name: `opts() has no duplicate keys (${keys.length} checked)`,
|
||||||
|
pass: dupes.length === 0, detail: [...new Set(dupes)].join(', ') });
|
||||||
|
} else {
|
||||||
|
out.push({ name: 'opts() literal found for duplicate-key check', pass: false, detail: '' });
|
||||||
|
}
|
||||||
|
|
||||||
|
const have = new Set([...html.matchAll(/id="([\w-]+)"/g)].map((m) => m[1]));
|
||||||
|
const missing = [...ids].filter((i) => !have.has(i));
|
||||||
|
out.push({ name: `every el() id exists in index.html (${ids.size} checked)`,
|
||||||
|
pass: missing.length === 0, detail: missing.join(', ') });
|
||||||
|
const unused = [...have].filter((i) => !ids.has(i));
|
||||||
|
out.push({ name: 'no orphaned ids in index.html', pass: unused.length === 0,
|
||||||
|
detail: unused.join(', ') });
|
||||||
|
} catch (e) {
|
||||||
|
out.push({ name: 'wiring check ran', pass: false, detail: e.message });
|
||||||
|
}
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
export function run() {
|
export function run() {
|
||||||
results.length = 0;
|
results.length = 0;
|
||||||
|
|
||||||
|
|
@ -113,6 +158,40 @@ export function run() {
|
||||||
const apRange = Math.max(...stab.aperture) - Math.min(...stab.aperture);
|
const apRange = Math.max(...stab.aperture) - Math.min(...stab.aperture);
|
||||||
ok('stabilisation preserves mouth motion', apRange > 0.05, `aperture range ${apRange.toFixed(4)}`);
|
ok('stabilisation preserves mouth motion', apRange > 0.05, `aperture range ${apRange.toFixed(4)}`);
|
||||||
|
|
||||||
|
// ASPECT: a shape that is circular in PIXEL space must stay circular in raster
|
||||||
|
// space. MediaPipe normalises x by width and y by height, so for a portrait
|
||||||
|
// frame equal normalised numbers are unequal pixel distances; feeding those
|
||||||
|
// straight through stretches everything horizontally by H/W. This asserts the
|
||||||
|
// isotropic conversion, and fails at ~1.78 for a 1080x1920 clip without it.
|
||||||
|
for (const [W, H] of [[1080, 1920], [1920, 1080], [640, 640]]) {
|
||||||
|
const aspect = W / H;
|
||||||
|
const N = 24, cx = 0.5, cy = 0.5, rPx = 200;
|
||||||
|
// a true circle of radius rPx, expressed in MediaPipe normalised coords
|
||||||
|
const circleFrames = [];
|
||||||
|
for (let t = 0; t < 4; t++) {
|
||||||
|
const pts = new Array(478).fill(null).map(() => ({ x: 0.5, y: 0.5, z: 0 }));
|
||||||
|
RIGID.forEach((id, k) => {
|
||||||
|
const a = (k / RIGID.length) * Math.PI * 2;
|
||||||
|
pts[id] = { x: cx + (120 * Math.cos(a)) / W, y: cy + (120 * Math.sin(a)) / H, z: 0 };
|
||||||
|
});
|
||||||
|
LIPS_OUTER.forEach((id, k) => {
|
||||||
|
const a = -(k / LIPS_OUTER.length) * Math.PI * 2;
|
||||||
|
pts[id] = { x: cx + (rPx * Math.cos(a)) / W, y: cy + (rPx * Math.sin(a)) / H, z: 0 };
|
||||||
|
});
|
||||||
|
FACE_OVAL.forEach((id, k) => {
|
||||||
|
const a = -(k / FACE_OVAL.length) * Math.PI * 2;
|
||||||
|
pts[id] = { x: cx + (420 * Math.cos(a)) / W, y: cy + (420 * Math.sin(a)) / H, z: 0 };
|
||||||
|
});
|
||||||
|
circleFrames.push(pts);
|
||||||
|
}
|
||||||
|
const st2 = stabilize(circleFrames, 0, aspect);
|
||||||
|
const ring = toRasterRing(st2.outer[0], LIPS_OUTER, 16, (p) => p);
|
||||||
|
const xs = ring.map((p) => p.x), ys = ring.map((p) => p.y);
|
||||||
|
const ratio = (Math.max(...xs) - Math.min(...xs)) / (Math.max(...ys) - Math.min(...ys));
|
||||||
|
ok(`circle stays circular at ${W}x${H}`, Math.abs(ratio - 1) < 0.02,
|
||||||
|
`w/h ratio ${ratio.toFixed(4)}`);
|
||||||
|
}
|
||||||
|
|
||||||
// key selection
|
// key selection
|
||||||
const xf = (p) => ({ x: p.x * 320, y: p.y * 200 });
|
const xf = (p) => ({ x: p.x * 320, y: p.y * 200 });
|
||||||
const shapes = stab.outer.map((r) => toRasterRing(r, LIPS_OUTER, 8, xf));
|
const shapes = stab.outer.map((r) => toRasterRing(r, LIPS_OUTER, 8, xf));
|
||||||
|
|
@ -127,6 +206,31 @@ export function run() {
|
||||||
ok('activeKey holds between keys',
|
ok('activeKey holds between keys',
|
||||||
activeKey(sel.keys, sel.keys[1].f - 1).f === sel.keys[0].f);
|
activeKey(sel.keys, sel.keys[1].f - 1).f === sel.keys[0].f);
|
||||||
|
|
||||||
|
// mouth lead: a shift that "feels like it does nothing" is indistinguishable
|
||||||
|
// from one that does nothing, so assert the arithmetic directly.
|
||||||
|
ok('lead 0 is identity', [0, 5, 71].every((f) => shiftIndex(f, 0, 72) === f));
|
||||||
|
ok('positive lead moves the source frame forward', shiftIndex(10, 2, 72) === 12);
|
||||||
|
ok('negative lead moves it back', shiftIndex(10, -3, 72) === 7);
|
||||||
|
ok('lead clamps at the start', shiftIndex(1, -6, 72) === 0);
|
||||||
|
ok('lead clamps at the end', shiftIndex(70, 6, 72) === 71);
|
||||||
|
{
|
||||||
|
// ...and that it selects different POSES, not merely different indices.
|
||||||
|
// Checked across the whole track rather than at one pair: synthetic poses
|
||||||
|
// hold for nine-frame beats, so any single pair can legitimately be
|
||||||
|
// identical while the shift works perfectly.
|
||||||
|
const N = shapes.length;
|
||||||
|
let moved = 0, total = 0;
|
||||||
|
for (let f = 0; f < N; f++) {
|
||||||
|
const a = shapes[shiftIndex(f, 0, N)], b = shapes[shiftIndex(f, 3, N)];
|
||||||
|
let d = 0;
|
||||||
|
for (let i = 0; i < a.length; i++) d += Math.hypot(a[i].x - b[i].x, a[i].y - b[i].y);
|
||||||
|
total++;
|
||||||
|
if (d / a.length > 0.5) moved++;
|
||||||
|
}
|
||||||
|
ok('a lead of 3 changes the pose on a good share of frames', moved / total > 0.2,
|
||||||
|
`${moved}/${total} frames differ`);
|
||||||
|
}
|
||||||
|
|
||||||
// rasteriser: indexed, hard-edged, no blending
|
// rasteriser: indexed, hard-edged, no blending
|
||||||
const r = new IndexedRaster(64, 48);
|
const r = new IndexedRaster(64, 48);
|
||||||
r.clear(0);
|
r.clear(0);
|
||||||
|
|
@ -147,6 +251,43 @@ export function run() {
|
||||||
ok('palette expansion introduces no intermediate colours',
|
ok('palette expansion introduces no intermediate colours',
|
||||||
[...seen].every((c) => allowed.has(c)), `${seen.size} distinct colours`);
|
[...seen].every((c) => allowed.has(c)), `${seen.size} distinct colours`);
|
||||||
|
|
||||||
|
// scaleRing is what pulls the sampled region in from MediaPipe's inner lip
|
||||||
|
// landmarks, which sit slightly outside the real opening.
|
||||||
|
{
|
||||||
|
const ring = [{ x: 0, y: 0 }, { x: 10, y: 0 }, { x: 10, y: 10 }, { x: 0, y: 10 }];
|
||||||
|
const small = scaleRing(ring, 0.5);
|
||||||
|
const w = Math.max(...small.map((p) => p.x)) - Math.min(...small.map((p) => p.x));
|
||||||
|
ok('scaleRing(0.5) halves the extent', Math.abs(w - 5) < 1e-9, `width ${w}`);
|
||||||
|
const same = scaleRing(ring, 1);
|
||||||
|
ok('scaleRing(1) is identity', same.every((p, i) => Math.abs(p.x - ring[i].x) < 1e-9));
|
||||||
|
let cx = 0; for (const p of small) cx += p.x;
|
||||||
|
ok('scaleRing keeps the centroid', Math.abs(cx / 4 - 5) < 1e-9);
|
||||||
|
}
|
||||||
|
|
||||||
|
// Otsu on a uniform region must report near-zero class separation. It will
|
||||||
|
// still return a threshold - that is what Otsu does - so the separation is the
|
||||||
|
// only thing that distinguishes "found teeth" from "split noise in a dark
|
||||||
|
// mouth", which is what made the band fill the whole cavity.
|
||||||
|
{
|
||||||
|
const flat = new Uint32Array(256); flat[40] = 500;
|
||||||
|
const f = otsuForTest(flat, 500);
|
||||||
|
ok('uniform region yields ~no class separation',
|
||||||
|
Math.abs(f.mBright - f.mDark) / 255 < 0.02, `sep ${((f.mBright - f.mDark) / 255).toFixed(4)}`);
|
||||||
|
|
||||||
|
const noisy = new Uint32Array(256);
|
||||||
|
for (let i = 30; i <= 60; i++) noisy[i] = 20; // dark cavity, some spread
|
||||||
|
const nz = otsuForTest(noisy, 31 * 20);
|
||||||
|
ok('dark-but-noisy region stays below a sane gate',
|
||||||
|
(nz.mBright - nz.mDark) / 255 < 0.14, `sep ${((nz.mBright - nz.mDark) / 255).toFixed(4)}`);
|
||||||
|
|
||||||
|
const teeth = new Uint32Array(256);
|
||||||
|
for (let i = 20; i <= 45; i++) teeth[i] = 40; // cavity
|
||||||
|
for (let i = 180; i <= 220; i++) teeth[i] = 30; // teeth
|
||||||
|
const tt = otsuForTest(teeth, 26 * 40 + 41 * 30);
|
||||||
|
ok('real bright/dark split clears the gate',
|
||||||
|
(tt.mBright - tt.mDark) / 255 > 0.4, `sep ${((tt.mBright - tt.mDark) / 255).toFixed(4)}`);
|
||||||
|
}
|
||||||
|
|
||||||
// take writer round-trip
|
// take writer round-trip
|
||||||
const take = {
|
const take = {
|
||||||
name: 'test', frames: 72, width: 320, height: 200, exposure: 2,
|
name: 'test', frames: 72, width: 320, height: 200, exposure: 2,
|
||||||
|
|
|
||||||
|
|
@ -1,4 +1,4 @@
|
||||||
// Take-file writer. Format is specified in docs/roto-puppet.md, "The take
|
// Take-file writer. Format is specified in docs/design.md, "The take
|
||||||
// format". Line-oriented text on purpose: Poco can parse it with fopen/fgets
|
// format". Line-oriented text on purpose: Poco can parse it with fopen/fgets
|
||||||
// from poco/src/safefile.c and strtok/atoi/atof from poco/src/strlib.c, so the
|
// from poco/src/safefile.c and strtok/atoi/atof from poco/src/strlib.c, so the
|
||||||
// Animator Pro render script needs no new native code.
|
// Animator Pro render script needs no new native code.
|
||||||
|
|
|
||||||
|
|
@ -10,11 +10,13 @@ import { applySim } from './mathutil.js';
|
||||||
|
|
||||||
// Compose pixel-space -> raster-space into one affine.
|
// Compose pixel-space -> raster-space into one affine.
|
||||||
//
|
//
|
||||||
// MediaPipe normalises x by width and y by height, so normalised space is a
|
// Landmarks are converted to an isotropic space (unit = one image height) before
|
||||||
// stretched pixel space and the composition is a general affine rather than a
|
// fitting, so pixels map in the same way: BOTH axes divide by imgH, not by their
|
||||||
// similarity. Three mapped points determine it exactly.
|
// own dimension. Dividing x by imgW here instead is what stretched the underlay
|
||||||
|
// horizontally by H/W and made it disagree with nothing - it matched the equally
|
||||||
|
// wrong vector shapes.
|
||||||
export function frameAffine(tf, xform, imgW, imgH) {
|
export function frameAffine(tf, xform, imgW, imgH) {
|
||||||
const map = (px, py) => xform(applySim(tf, { x: px / imgW, y: py / imgH }));
|
const map = (px, py) => xform(applySim(tf, { x: px / imgH, y: py / imgH }));
|
||||||
const P0 = map(0, 0), P1 = map(imgW, 0), P2 = map(0, imgH);
|
const P0 = map(0, 0), P1 = map(imgW, 0), P2 = map(0, imgH);
|
||||||
return {
|
return {
|
||||||
a: (P1.x - P0.x) / imgW, b: (P1.y - P0.y) / imgW,
|
a: (P1.x - P0.x) / imgW, b: (P1.y - P0.y) / imgW,
|
||||||
|
|
|
||||||
|
|
@ -1 +1 @@
|
||||||
{"fps":12,"frames":37,"dir":"frames","audio":"audio.wav","source":"IMG_8486.MOV"}
|
{"fps":24,"frames":74,"dir":"frames","audio":"audio.wav","source":"IMG_8486.MOV"}
|
||||||
|
|
|
||||||
79
probe.html
79
probe.html
|
|
@ -1,79 +0,0 @@
|
||||||
<!doctype html><html><head><meta charset="utf-8"><title>probe…</title></head>
|
|
||||||
<body><pre id="o">running…</pre>
|
|
||||||
<script type="module">
|
|
||||||
import { FaceLandmarker, FilesetResolver } from 'https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@1.0.1/vision_bundle.mjs';
|
|
||||||
import { LIPS_OUTER, LIPS_INNER, subsampleSlots } from './js/landmarks.js';
|
|
||||||
import { stabilize, toRasterRing, selectKeys } from './js/pipeline.js';
|
|
||||||
|
|
||||||
const log = [];
|
|
||||||
const say = (s) => { log.push(s); document.getElementById('o').textContent = log.join('\n'); };
|
|
||||||
|
|
||||||
function crosses(a,b,c,d){const o=(p,q,r)=>Math.sign((q.x-p.x)*(r.y-p.y)-(q.y-p.y)*(r.x-p.x));
|
|
||||||
const o1=o(a,b,c),o2=o(a,b,d),o3=o(c,d,a),o4=o(c,d,b);
|
|
||||||
return o1!==o2&&o3!==o4&&o1!==0&&o2!==0&&o3!==0&&o4!==0;}
|
|
||||||
function selfInts(pts){const n=pts.length,h=[];for(let i=0;i<n;i++)for(let j=i+1;j<n;j++){
|
|
||||||
if((j+1)%n===i||(i+1)%n===j)continue;
|
|
||||||
if(crosses(pts[i],pts[(i+1)%n],pts[j],pts[(j+1)%n]))h.push([i,j]);}return h;}
|
|
||||||
|
|
||||||
const loadImg = (src) => new Promise(r => { const i=new Image(); i.onload=()=>r(i); i.onerror=()=>r(null); i.src=src; });
|
|
||||||
|
|
||||||
try {
|
|
||||||
const imgs=[];
|
|
||||||
for(let i=1;i<=900;i++){const im=await loadImg(`frames/${String(i).padStart(4,'0')}.png`); if(!im)break; imgs.push(im);}
|
|
||||||
say(`frames loaded: ${imgs.length} @ ${imgs[0].naturalWidth}x${imgs[0].naturalHeight}`);
|
|
||||||
|
|
||||||
const fs = await FilesetResolver.forVisionTasks('https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@1.0.1/wasm');
|
|
||||||
const lm = await FaceLandmarker.createFromOptions(fs, {
|
|
||||||
baseOptions:{ modelAssetPath:'./face_landmarker.task', delegate:'CPU' },
|
|
||||||
runningMode:'IMAGE', numFaces:1 });
|
|
||||||
say('landmarker ready (CPU delegate)');
|
|
||||||
|
|
||||||
const cv=document.createElement('canvas'); const dense=[]; let miss=0;
|
|
||||||
for(const im of imgs){
|
|
||||||
cv.width=im.naturalWidth; cv.height=im.naturalHeight;
|
|
||||||
cv.getContext('2d').drawImage(im,0,0);
|
|
||||||
const out=lm.detect(cv);
|
|
||||||
if(out.faceLandmarks?.length) dense.push(out.faceLandmarks[0]);
|
|
||||||
else { miss++; if(dense.length) dense.push(dense[dense.length-1]); }
|
|
||||||
}
|
|
||||||
say(`detected: ${dense.length}/${imgs.length} (no face on ${miss})`);
|
|
||||||
if(!dense.length) throw new Error('no face detected in any frame');
|
|
||||||
|
|
||||||
// THE key check: are LIPS_OUTER / LIPS_INNER correct traversals of real data?
|
|
||||||
for(const [name,tab] of [['LIPS_OUTER',LIPS_OUTER],['LIPS_INNER',LIPS_INNER]]){
|
|
||||||
let bad=0, first=null;
|
|
||||||
for(let n=4;n<=16;n+=2){
|
|
||||||
const slots=subsampleSlots(tab.length,n);
|
|
||||||
for(let f=0;f<dense.length;f++){
|
|
||||||
const h=selfInts(slots.map(s=>dense[f][tab[s]]));
|
|
||||||
if(h.length){bad++; first=first||`verts=${n} f=${f} edges ${JSON.stringify(h[0])}`;}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
say(`${name}: ${bad===0?'SIMPLE at every budget/frame':`SELF-INTERSECTS ${bad}x first ${first}`}`);
|
|
||||||
// full 20-ring too
|
|
||||||
let bad20=0;
|
|
||||||
for(let f=0;f<dense.length;f++) if(selfInts(tab.map(i=>dense[f][i])).length) bad20++;
|
|
||||||
say(` full 20-point ring: ${bad20===0?'simple on all frames':`self-intersects on ${bad20} frames`}`);
|
|
||||||
}
|
|
||||||
|
|
||||||
const st = stabilize(dense, 5);
|
|
||||||
const res = st.residual;
|
|
||||||
const mean = res.reduce((a,b)=>a+b,0)/res.length;
|
|
||||||
say(`residual mean ${mean.toFixed(5)} max ${Math.max(...res).toFixed(5)} (high = out-of-plane rotation)`);
|
|
||||||
const eyeX = st.eyes.map(e=>e[0].x);
|
|
||||||
const rawX = dense.map(f=>f[133].x);
|
|
||||||
say(`eye-inner x spread: raw ${(Math.max(...rawX)-Math.min(...rawX)).toFixed(4)} -> stabilised ${(Math.max(...eyeX)-Math.min(...eyeX)).toFixed(4)}`);
|
|
||||||
const ap=st.aperture;
|
|
||||||
say(`aperture min ${Math.min(...ap).toFixed(4)} max ${Math.max(...ap).toFixed(4)} range ${(Math.max(...ap)-Math.min(...ap)).toFixed(4)}`);
|
|
||||||
let nf=0; ap.forEach((v,i)=>{ if(v===Math.min(...ap)) nf=i; });
|
|
||||||
say(`most-closed frame: ${nf}`);
|
|
||||||
|
|
||||||
const xf=p=>({x:p.x*320,y:p.y*200});
|
|
||||||
const shapes=st.outer.map(r=>toRasterRing(r,LIPS_OUTER,8,xf));
|
|
||||||
for(const [mh,dt] of [[1,0.6],[2,0.6],[2,1.5],[2,3.0],[3,1.5]]){
|
|
||||||
const k=selectKeys(shapes,{minHold:mh,distThresh:dt,velSmooth:3,exposure:1});
|
|
||||||
say(`minHold=${mh} gate=${dt}: ${k.candidates.length} cand -> ${k.keys.length} keys [${k.keys.map(x=>x.f).join(' ')}]`);
|
|
||||||
}
|
|
||||||
document.title='PROBE OK';
|
|
||||||
} catch(e){ say('ERROR: '+e.message+'\n'+e.stack); document.title='PROBE FAIL'; }
|
|
||||||
</script></body></html>
|
|
||||||
|
|
@ -7,10 +7,11 @@
|
||||||
</style></head><body>
|
</style></head><body>
|
||||||
<h1 id="head">running…</h1><ul id="out"></ul>
|
<h1 id="head">running…</h1><ul id="out"></ul>
|
||||||
<script type="module">
|
<script type="module">
|
||||||
import { run } from './js/selftest.js';
|
import { run, runWiring } from './js/selftest.js';
|
||||||
let res;
|
let res;
|
||||||
try { res = run(); }
|
try { res = run(); }
|
||||||
catch (e) { res = [{ name: 'harness threw: ' + e.message, pass: false, detail: String(e.stack).split('\n')[1] || '' }]; }
|
catch (e) { res = [{ name: 'harness threw: ' + e.message, pass: false, detail: String(e.stack).split('\n')[1] || '' }]; }
|
||||||
|
res = res.concat(await runWiring());
|
||||||
const pass = res.filter(r => r.pass).length;
|
const pass = res.filter(r => r.pass).length;
|
||||||
const fail = res.length - pass;
|
const fail = res.length - pass;
|
||||||
document.getElementById('head').textContent = `${fail === 0 ? 'PASS' : 'FAIL'} — ${pass}/${res.length} assertions`;
|
document.getElementById('head').textContent = `${fail === 0 ? 'PASS' : 'FAIL'} — ${pass}/${res.length} assertions`;
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue