Commit graph

14 commits

Author SHA1 Message Date
Your Name
c2603102da readme additions 2026-09-24 15:49:05 -04:00
f827757f7d Become arthur: a standalone suite, not an Animator Pro front-end
The test renderer turned out to be the product. Everything that decides how the
work looks - stabilisation, reduction, timing, frame removal, palette - already
happens here, and the flat indexed output already reads the way it should.

The reason to leave is in the original design's own rule: never make a timing
decision that requires a full render to evaluate. Honouring that moved every
judgement out of Animator Pro, which left the host doing nothing but writing a
file, in exchange for modal UI, minutes-long renders, one-level undo, FLX delta
invariants, a single tween state and a cel singleton.

What does NOT change is the constraint. 320x200, indexed palette, flat fills,
no antialiasing - inherited, but load-bearing rather than accidental. The
rasteriser writes palette indices and expands to RGBA only at the end precisely
so nothing can soften an edge. Modern conveniences belong in the workflow.

Adds docs/design.md: the principles, carried over without the Poco/FLX/cel
machinery, plus architecture and an honest list of what is missing - the
largest gap being that plates still have nowhere to be drawn.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:47:41 -04:00
dc0a3f34e6 Make the mouth lead legible and assertable
"It feels like it does nothing" is indistinguishable from "it does nothing",
and at 24fps a lead of 1 is 42ms - small enough to reasonably doubt. So the
shift is now provable and visible rather than taken on trust:

- shiftIndex is a pure exported function with assertions covering identity,
  both directions, and clamping at each end
- the frame label always shows the mouth frame, not only when shifted, so the
  number can be watched diverging from f; non-zero leads also report in ms
- the slider readout carries an explicit + sign
- the stabilised pane draws the unshifted contour as a dark-green ghost when a
  lead is set, so the offset is something you can see

Also: leadIndex called opts() on every invocation - ~14 DOM reads, once per
strip thumbnail, so ~1000 per redraw at 74 frames. It reads a cached scalar now.

The "different pose" assertion checks the whole track rather than one pair:
synthetic poses hold for nine-frame beats, so a single pair can legitimately be
identical while the shift works correctly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:38:02 -04:00
f9b8ec8617 Fix teeth vertex slider driving the lip vertex count
opts() declared `verts` twice - once from the lip slider and again from the
teeth slider. Duplicate keys in an object literal are silent in JS and the last
one wins, so the teeth vertex control was quietly setting the lip vertex budget
while the lip control did nothing at all.

Renamed to teethVerts, with interior.js reading it under that name.

selftest now parses the opts() literal and fails on duplicate keys, verified to
catch this exact case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:36:04 -04:00
d18791e6ad Mouth lead: nudge performance tracks ahead of the clock
Smoothing does not literally delay anything - a centred moving average has zero
phase lag - but it blurs onsets, so the visually salient moment of a mouth
opening moves later even though the mean does not. Animators also draw mouth
shapes a frame or two ahead of the sound as standard practice, so this is the
normal control rather than a workaround.

Only the performance tracks shift; the head and audio stay put, since it is the
mouth that should anticipate. In photo-underlay modes the vector mouth will
therefore no longer match the frame behind it, which is expected.

The lead is baked into the exported take - key f carries the pose from source
frame f+lead - so the Animator Pro renderer never needs to know about it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:34:08 -04:00
a146d271b6 Cross-check el() ids against index.html in the selftest
A knob wired in app.js but missing from the markup throws during wiring, aborts
the module and leaves a blank page - a symptom pointing nowhere near its cause.
It has now happened twice, so it gets a check rather than vigilance: the
selftest fetches both files and compares the id sets in each direction.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:25:39 -04:00
9a11eeabc1 Teeth as an extracted blob contour, not a clipped band
The band filled the mouth because a band is the wrong reduction: the bright
region is a blob, and reading it as "everything above a line" throws the shape
away.

Extracting a contour reintroduces the vertex-correspondence problem that made
me avoid it, but for a blob there is a way out. Radial sampling from the
centroid along N fixed directions makes vertex k always mean "the extent in
direction k": correspondence holds by construction, the count is fixed, and
temporal smoothing cannot reorder anything. It also yields a star-shaped
reduction, which suits flat colour.

Tongue rejection, which the band had no way to express:
- pixels red relative to their own brightness are dropped (teeth are neutral)
- component choice is biased toward the top of the cavity, since area alone
  picks the tongue when the mouth is wide
- separate inner and outer controls: cavity erode pulls the sampled region off
  the lip edge, blob grow/erode resizes the found blob

Also: a knob wired in app.js but missing from index.html threw during wiring
and left a blank page with nothing useful in the console - which is exactly
what happened to teethDwell in the previous commit. el() now names the missing
id, and window.onerror surfaces it in the status line.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:24:52 -04:00
c6daf827d1 Fix teeth band filling the whole mouth
Three causes, all of them mine:

Otsu always returns a split, including on a homogeneous region - given a dark
cavity with no teeth it invents a threshold and calls half the pixels bright.
The gate is now the separation between the two class means, which is the only
thing that says whether the split means anything. Coverage was the wrong
signal: it is high both when the mouth is full of teeth and when the region is
uniformly dark and Otsu has split noise.

The row scan tracked the last qualifying row anywhere rather than where the run
from the top stops, so one bright row near the bottom - a lit lower lip inside
the ring - pushed the line to full height. It now breaks at the first failing
row once the run has started.

MediaPipe's inner lip landmarks sit slightly outside the real opening, so the
sampled region included lip pixels, which are bright and sit exactly at the
boundary where they do most damage. The ring is now eroded toward its centroid
before sampling, with the amount exposed as a knob.

Adds a diagnostic panel showing the sampled crop, pixels above threshold, and
the resolved line, because tuning this from numbers alone does not tell you
whether the region being measured is even the right region.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:14:49 -04:00
941022b69f Teeth from mouth-interior image content
MediaPipe has no landmarks inside the lips, so teeth have to come from pixels.
Tracing the bright blob would give a new contour every frame with no vertex
correspondence - the exact boil docs/roto-puppet.md warns about.

So the measurement yields a scalar, not a shape: Otsu within the cavity, scanned
from the top for where the bright run stops, giving one line height per frame.
The teeth polygon is the inner lip ring clipped to that line, so the silhouette
is always the mouth's own shape and cannot disagree with the lips around it,
and the only per-frame variable is a single number that smooths trivially.

Presence gets hysteresis and minimum dwell, as plate selection does: a teeth
block blinking on and off for single frames is worse than one simply absent.

Tongue is not implemented. The same scalar approach would apply, gated on
redness rather than brightness, but it is not visible in the test footage - the
cavity reads dark with a bright upper-teeth band and nothing else.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:12:09 -04:00
29b0bd6690 Fix horizontal stretch from MediaPipe's anisotropic normalised space
MediaPipe normalises x by image WIDTH and y by image HEIGHT, so for a 1080x1920
clip one unit of x is 1080px and one unit of y is 1920px. makeXform applied a
single scale to both, stretching everything horizontally by H/W - 1.78x on this
footage. The photo underlay looked equally squashed because frameAffine divided
x by imgW, matching the equally wrong vector shapes rather than disagreeing
with them.

Fixed at ingest: landmarks convert to an isotropic space whose unit is one image
height (x *= W/H), so equal numbers mean equal pixels everywhere downstream.
Pixel mapping follows - both axes divide by imgH.

This also silently fixes head roll. fitSimilarity was fitting a rotation in a
sheared space, so the "similarity" it recovered was not one, and stabilisation
of rolled heads was subtly wrong.

selftest: a shape circular in pixel space must stay circular in raster space,
checked at 1080x1920, 1920x1080 and 640x640. Fails at ratio 1.78 without the
conversion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:03:12 -04:00
de90bfd907 Registered photo underlay as the plate reference
The generated face oval was never going to be good enough to draw from:
MediaPipe's face oval is the FACE boundary, cut at the hairline and excluding
hair, ears, jaw and neck, so it is an egg by construction. Segmentation would
give a real head outline but costs a 16MB model and per-frame inference for a
shape that gets replaced by a drawing anyway.

So the plate layer becomes switchable, and the useful modes are photographic:
the source frame mapped into raster space through the same transform chain the
contours go through. Registration is the whole point - the head sits still and
a drawing traced from the underlay is already aligned to the mouth. An
unregistered underlay would be decoration.

- underlay.js: pixel->raster affine (a general affine, since MediaPipe
  normalises x by width and y by height), registered draw, palette posterise
- plate modes: photo / photo dim / posterized / oval / oval+photo / none, B cycles
- worksheet cells are registered composites rather than raw crops
- Save frame 4x writes a 1280x800 PNG to draw on
- selftest: FACE_OVAL simplicity, which was never asserted; a wrong ordering
  there reads as a lumpy plate rather than an obvious bowtie

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 14:57:59 -04:00
5ac919b9d6 Audio-clocked playback for judging sync
Lip sync cannot be judged silently. extract.sh now pulls the audio track and
writes manifest.json alongside the frames; the page reads the true extraction
rate from it rather than assuming one, since a guessed fps desynchronises
picture from sound - the one thing this view exists to show.

Audio is the clock: frame = floor(currentTime * fps). A slow render loop drops
frames instead of drifting, and half/quarter speed work via playbackRate with
the picture following for free. Scrubbing, stepping and clicking a thumbnail
all seek the audio too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 14:53:09 -04:00
a082208ad5 Frame removal for plates; mouth keeps every frame
Sparseness was being applied for two different reasons at once. Aesthetic
sparseness is set by the extraction rate; labour sparseness only binds on the
plate, because a human draws each one. The mouth is traced and therefore free,
and in limited animation lip sync is routinely the densest element - on 1s
while heads hold on 2s and 3s.

So: the mouth gets a key on every frame, and the frame strip is now the
editing surface for deciding which frames need their own plate drawing. All
frames start kept; delete the ones you don't want.

- strip of face-cropped thumbnails, keep/drop per frame, keyboard driven
- worksheet panel lists the drawings needed and the range each one holds
- Suggest runs error-tolerance decimation on head pose as a starting point
- export writes sparse plate keys + dense mouth keys, with a hold manifest
- smoothContours: bounded exception to "never smooth the contour", which held
  only while keys were sparse enough to reject detector noise by sampling
- averages are now a RADIUS in frames: 0 is off, 1 is +-1
- GPU delegate falls back to CPU instead of failing
- #synth / #frames autorun for headless smoke tests

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 14:51:15 -04:00
7a6bdea157 roto: video -> take file builder with interactive tuning
Analysis half of the pipeline in docs/roto-puppet.md. Stabilises a face out
of a clip via a similarity fit on rigid landmarks, reduces the lip contour to
a fixed vertex budget, selects sparse keys on velocity minima, and previews
the result as flat indexed fills so timing can be judged without an Animator
Pro render.

- landmarks.js  ordered lip/oval rings; slot position is vertex identity
- mathutil.js   closed-form 2D similarity, Procrustes mean, transform smoothing
- pipeline.js   stabilise -> subsample -> key-select
- raster.js     indexed scanline fill, no antialiasing
- take.js       take-file writer
- selftest.js   29 assertions, incl. ring simplicity at every vertex budget

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 14:38:07 -04:00