Commit graph

125 commits

Author SHA1 Message Date
Olive Vaughn
8a06835895 Port step 5: freeze measured mouth into playable channels 2026-09-27 19:01:39 -04:00
Olive Vaughn
942e2f38ab Port step 4: measure the anchor and the mouth, condition on its own
`stabilize` is three things wearing one name, and it is now three functions in two
stages: `flow/measure/anchor` fits the rigid transform, `flow/condition` smooths
its parameters, `flow/measure/mouth` measures the lip rings through the result.
Parity is on the COMPOSITION and not on the pieces -- a split that agreed
function by function and not end to end would be a split rather than a port.

The oracle now drives `stabilize` at three configurations and the port agrees to
1e-9 on ref, rigid, transforms, outer, inner and aperture, plus `smoothContours`
at three radii. Two of the three configurations are at aspect 0.5625, a 1080x1920
phone clip, because at aspect 1 `pick` is the identity: a port that dropped the
anisotropy correction outright would pass every other assertion in the suite.
148 tests, up from 134.

Three decisions worth the reading time.

`makeXform` is not ported, and its absence takes the face oval with it. It
centres on the oval's bounding box and zooms until the face is 80% of the raster
height, so every vertex it touched carried a cropping decision made once, at
analysis time, from one frame's landmarks. Geometry belongs in the node's own
local space with the framing as a transform on a node, so this is a deletion. The
oval's only other consumer was the placeholder plate outline, which is painting.

The residual is taken against the RAW fit, and the prototype took it against the
smoothed one. That is the only deliberate numeric divergence here, and parity is
kept by asserting `anchor/residuals` on exactly what the prototype handed it. The
number's job is to say whether a section is stabilisable at all; folding the
smoothing error into it makes a slider look like a property of the footage, and
docs/architecture.md lists the residual under stage 3, which requires it to be
knob-free. `condition/anchor` therefore replaces `:transforms` and leaves
`:residual` alone.

The stage order is not the strict chain the table in docs/architecture.md looks
like, and that document now says so. The fit is knob-free, conditioning smooths
it, and the rings are measured *through* the conditioned transform -- so
`anchor avg` does re-run the ring mapping, which is a few hundred frames of twenty
points. The guarantee was only ever about the part that reads a source pixel, and
that part never sees a transform.

Two things fall out and are asserted rather than assumed. Smoothing and
subsampling commute, because both are per-slot, which is what lets `vertices`
stay a stage-5 knob downstream of a stage-4 one -- and it is also why the port can
smooth the full twenty slots where the prototype smooths eight and still match.
And `condition/contours` is `geom/moving-average` per vertex per axis rather than
its own clamped window, so "radius 2" cannot come to mean two different things at
the two knobs.

One dead end recorded so nobody walks it twice: the synth's head is perfectly
rigid -- its jitter is a whole-head translation, which a similarity absorbs
exactly -- so every frame's rigid configuration is congruent with frame zero's and
the Procrustes mean IS frame zero to 1e-15, jitter or none. "The reference is the
mean and not frame zero" cannot be asserted on this track and is asserted in
geom-test, where the two can differ.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-27 18:00:11 -04:00
Olive Vaughn
11192d61c6 Speed up the frame, and drop loop/recur from the domain
The frame went 5.86ms to 3.17ms -- a 170fps ceiling to 315 -- and `loop`/`recur`
is gone from src/ entirely.

Two real wins, both from measuring rather than guessing:

- `->rgba` was 3.16ms of that frame and was scene-independent: a `nth` into a
  vector of vectors is four protocol dispatches per pixel, 64,000 pixels a frame.
  The palette is now flattened once and cached by identity of the source vector
  -- palettes are values, so identity is exactly the right test and there is no
  invalidation to get wrong. At zoom 1 on a little-endian machine the inner loop
  is one 32-bit write per pixel through a Uint32Array view of the same buffer:
  0.11ms, 28x. Every other case walks bytes off the same flat palette.
  raster-test pins both against a naive per-pixel reference at three zooms,
  because a fast path that is subtly wrong about colour would look like a palette
  bug rather than like an optimisation.

- The per-frame z sort was re-deriving a constant. Draw order is a function of
  the z paths, which change when the scene changes and never because the playhead
  moved, so `draw-rank` computes it once and a frame sorts small integers. Every
  op drops its `:i` and `:z-path` fields as a result.

The loop pass, and an honest note on it: it came out NET POSITIVE on lines, which
is the wrong direction for a cleanup. geom is -3 (transduce for the accumulators,
`(-> (iterate refine ref) (nth iters))` for Procrustes, which is what the
algorithm says rather than a counter that happens to stop), channel -3,
fill-poly!'s copy loop 7 lines to 1. Against that, eval-into went from one
four-deep pyramid with seven positional parameters to `place` / `emit` / a fold
over a ctx map -- less nesting, more lines, and a different change from "fix the
loops" that should not have been bundled with it.

Two idioms were reverted for being worse here than what they replaced, both the
same mistake -- reaching for a form that allocates inside a hot loop:

- `partition 2` over an `array-seq` per scanline is some five thousand throwaway
  objects a frame and took draw from 0.88ms to 1.48ms. Now a pairwise `dotimes`
  over the array.
- `z-lex` via `(map compare a b)` allocated three lazy seqs per call, ~700 calls
  a frame. Made moot by `draw-rank`.

And one DRY move reverted for coupling things that only coincide: a `geom-path`
table had `node/valid-paths` and `scene/emit` deriving from one source, which
ties what a kind may CARRY to what the renderer READS off it. Those are the same
today and are not the same question, and the table put a spec change in charge of
what gets drawn, across a namespace boundary. `emit`'s three branches are three
different marks and stay three branches.

Kept, because it is one operation with two callers rather than two concerns that
rhyme: `lineage`, which `depth` and `z-path` were both walking separately. Its
cycle check is now a length bound -- a chain that does not repeat cannot be
longer than the node count -- instead of a `seen` set.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PDfHGdV39zu6rvgbBTfDaT
2026-09-27 17:44:07 -04:00
Olive Vaughn
18d6495592 Port steps 2-3: the data model and the player
Steps 2 and 3 land together because the model revisions in the middle changed
code from both, and splitting them now would invent intermediate states that
never built.

  domain/channel  value-at across framed/keyed/dense, plus a cursor
  domain/node     decomposed transform, composition order, time maps
  domain/scene    topological order, z paths, eval-frame and resolver
  clock           audio-clocked frame derivation, outside app-db
  db/events/subs  re-frame arrives; the playhead is document state
  ui/player       the rAF loop; reads, blits, dispatches (almost) nothing
  ui/shell        transport

133 tests, 1158 assertions. The scene plays at 30fps against audio, scrubs, and
runs at 1/4x through 4x; verified by driving a real browser over CDP rather than
by assertion.

Two evaluators, on purpose. `eval-frame` is the specification -- allocating,
order-free, obviously correct. `resolver` is what playback uses: cached topo
order and z paths, a cursor per channel, a preallocated point buffer per node.
Both run the same walk, parameterised only by how a channel is read and where
points are written, because two independent implementations of frame evaluation
would drift and the drift would read as a rendering bug rather than as two
functions disagreeing. scene-test asserts they agree frame for frame in forward,
backward and random order.

Deviations and decisions, each with a reason:

- raster/fill-poly! is now a thin wrapper over fill-poly-buf!, which takes a flat
  preallocated buffer. ONE scanline fill serves the analysis stages, which speak
  {:x :y}, and frame evaluation, which hands over a buffer it owns. The parity
  suite still passes pixel-for-pixel, which is what makes the rewrite safe.

- The state mask carries ABSENCE ONLY. An earlier draft gave it a hidden bit too,
  per architecture.md's "hidden flag + palette index", and that bit was a dense
  [:vis] wearing a different hat -- two mechanisms for one question, which is how
  a part ends up hidden by one and shown by the other.

- The palette is a parameter of evaluation, not a global. A node names a TONE;
  which ramp that tone is read in belongs to the timeline it sits in.

- :over layers and a symbol :rate THROW rather than being ignored. Neither is
  built and nothing can produce one, so this can only fire on data that has run
  ahead of the code. A silently dropped override is a hand correction the user
  made once, watched fail, and has no reason to trust again.

Three findings the model produced rather than received:

- Presence propagates asymmetrically. An absent transform drops the subtree; an
  absent [:geom :pts] drops only that node, because an absent mouth outline has
  nothing to draw but the head it hangs off has not moved. That asymmetry is the
  reason presence is tracked per channel and not per node.

- Z paths need lexicographic compare, not `compare`, which orders vectors by
  count first -- so a cel three levels under "a1" would jump in front of a bare
  "a2" and the layer order would mostly work.

- A node stencilled by something that drew nothing is dropped, not drawn
  unclipped: an iris floating over the cheek is worse than a missing iris.

docs/ revised alongside, and those revisions are the load-bearing part:

- A scene, a timeline and a symbol are one type. The doc had two structures with
  the same fields and never said so. Two axes of nesting are now separated --
  parent/child within a timeline is flat with parent pointers, instance nesting
  is by reference -- which is why "nestable" and "flat" only sounded
  contradictory.

- Palettes are named, live on the project, and are ENABLED on a timeline as a
  channel. Absent inherits; present travels with the timeline, so a symbol
  authored against :night stays night wherever it is placed. The output index
  space is the concatenation of the named ramps, which keeps one buffer and one
  flat table and incidentally stops two nodes in different palettes colliding on
  a stencil.

- Stabilisation is a channel, not a mode: {s, theta, tx, ty} IS [:xform :*], so
  the normalise on/off/per-plate toggle is which of the three channel shapes the
  :head node carries. Always measure and always store factored -- smoothing and
  velocity-minimum key selection both need the split to exist in storage.

- There is no camera node and none is needed. Placement is a node transform, the
  stage clips what hangs off it, and project dimensions are independent of the
  footage. `makeXform` is therefore not to be ported: it bakes a cropping
  decision into every stored vertex.

- Export is removed. The .take writer was for an Animator Pro render script; the
  target is encoding video in the browser, and step 9 now says not to port the
  old one.

demo/swarm is 120 shapes on six orbits, entirely dense blocks behind store
handles -- the shape freeze produces at step 5, and the first thing to exercise
that path under load. It plays at 30fps, and bench-test keeps a deliberately
loose floor under it because a performance regression here does not announce
itself: the picture stays correct and merely arrives late.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PDfHGdV39zu6rvgbBTfDaT
2026-09-27 17:28:05 -04:00
Olive Vaughn
eb06be005c Port steps 0-1: scaffold, the oracle, and the pure bottom
Scaffolds frontend/ (shadow-cljs, reagent 1.2.0, re-frame 1.4.3) and ports
everything below the data model, with the JS kept as a numeric oracle.

  domain/landmarks  index tables, verbatim
  domain/ring       subsample, offset, simplicity
  domain/geom       similarity fit, procrustes, moving average
  domain/raster     indexed scanline fill, stencil, disc, rect
  domain/palette    the ramp, and the no-sampled-RGB rule

58 tests, 166 assertions. Parity with js/ on the identical 72-frame synthetic
track: fit-similarity, procrustes-mean, fit-residual, moving-average,
smooth-transforms, offset-ring and subsample-slots to 1e-9; the raster
pixel-for-pixel over the whole buffer.

Three deviations from the JS, each for a reason:

- synth.cljs jitters from a SEEDED generator, not Math.random. Parity is only
  checkable if both sides can be handed the same track, and a failing assertion
  has to be reproducible. `:rand-fn` takes the generator over, so oracle.mjs
  stubs js/Math.random and js/ itself stays untouched.

- raster/->rgba replaces toImageData. ImageData is a DOM type and domain/ may
  not touch the DOM; returning plain bytes also lets the
  no-intermediate-colours assertion run in node. ui/canvas wraps it later.

- offset-ring lives in domain/ring, not domain/geom, per architecture.md: it is
  an operation on an ordered traversal, not on a transform.

Step 1's "done" also names the swapped-iris vote, but pairIrises is in
pipeline.js and belongs to step 7. The precondition is asserted instead --
`:swap-iris` really does move both blocks -- so the vote will have a track that
disagrees with it when it arrives.

One finding, recorded in full in the test that measures it: smooth-transforms
buys nothing on the synthetic track. Against jitter-free ground truth, radius 1
helps by 17% on one noise realisation and hurts by 0.5% on another, so its
benefit is within noise; from radius 2 up the cost is unambiguous, and by radius
5 the filter is below the true motion's own high-frequency energy, i.e.
smoothing away performance. The test pins the shape of the knob rather than a
preferred value. This may say more about the synth's jitter being unrealistically
small (+/-0.001 normalised) than about the knob; step 6 settles it on real
footage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-27 14:43:34 -04:00
Your Name
082d8561d2 idk 2026-09-25 12:55:01 -04:00
Your Name
b68deb838a Paint: copy previous should skip empty frames
Copy previous looked for the nearest earlier ENABLED frame, and every frame is
enabled until you thin the strip out - so it resolved to f-1, which has nothing
on it, and the button appeared to only ever copy the immediate sibling while
reporting "nothing to copy".

It now takes the nearest earlier enabled frame that actually carries a drawing.
That behaves identically before and after you curate the strip, which is the
point: the rhythm of the drawings should not depend on whether you have got
round to deleting frames yet.

Cels stay tied to the keep-set. They are plate drawings and they hold until the
next enabled frame, as originally specified - an earlier version of this commit
gave them their own independent set, which is wrong for what they are.

Frames carrying a drawing are now marked in the strip, because "copy previous"
reaching back to a frame you cannot see is not much better than it reaching to
the wrong one.

Also fixes stale paint labels: onChange refreshed the strip and the panes but
not the paint header, so the drawn-frames readout lagged a copy behind. Split
the text update out of drawPaint so it can run from inside commit() without
re-rendering the canvas underneath itself.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 19:25:15 -04:00
Your Name
9d516a03ac Paint: vector background cels on the kept frames
A sketch, and labelled as one. It exists to test whether the aesthetic holds
when a human draws the background rather than the tracker deriving it, and it
is meant to be replaced by a real paint surface with onion skin and undo. Kept
to one dependency-free module so throwing it away is a delete, not surgery.

Cels are drawn on the frames that get their own drawing and hold until the
next one - the same rule the plate follows, and literally the same lookup, so
the two cannot disagree about what is on screen. Editing always targets the
cel you can see, so you can scrub anywhere and keep drawing.

Pen places vertices and closes on the first one. Edit drags vertices or whole
shapes, shift-click inserts, alt-click removes. Layers stack front-at-top with
per-layer colour, show/hide and reorder. Drag a strip thumbnail onto the canvas
to seed this cel from that one, every layer, as a deep copy - sharing the point
arrays would make two cels silently edit each other.

Two rules are enforced rather than left to discipline:

Colours are PALETTE INDICES, never RGB. Sampling colour from the source is the
one move docs/design.md calls irrecoverable, and a paint tool is exactly where
that discipline would leak, so the picker cannot express a colour outside the
ramp.

Vertices snap to the 320x200 grid. On a hard-edged indexed rasteriser a shape
nudged by 0.4px moves an edge by a whole pixel or not at all depending on where
it happens to land, so sub-pixel vertices shimmer instead of holding still.

Drawings autosave to localStorage per take name. They are the only thing here a
person made by hand; everything else regenerates. Not in the .take export yet.

Also adds serve.py, a no-store dev server. python3 -m http.server sends
Last-Modified and browsers cache ES modules on it hard enough that a reload
serves a stale app.js against a fresh index.html: the new knobs appear, nothing
wires them, no error fires, and it reads as "your feature does not work". That
cost real time this session.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 19:08:07 -04:00
Your Name
03df81fa0b Brows: traced ring, quantised raise
A brow at 320x200 is fourteen pixels wide and three tall. Its shape carries
almost nothing at that size; its height above the eye carries the expression,
and a brow raise is the most legible beat on a face. So the ring is traced and
the height is quantised - the split the eyes already got, where the lid is a
traced feature and the iris a quantised primitive.

The decomposition is the point. The traced ring already contains the real
height, so adding a quantised raise on top would move the brow twice. The
height is measured OUT of the ring, quantised, and put back, so the shape that
renders is his at a height that snaps between a few levels and holds.

Measured at both ends rather than as one number, because raise and tilt are
different expressions out of one mechanism: both ends up is surprise, inner up
alone is worry, inner down is anger. They share a dwell - the gaze quantiser,
renamed quantizeSnap now that it has two callers - so the brow hits its pose in
one frame instead of crawling into it with one end arriving before the other.

Measured against the eye's corner midpoint, never its lid. Same trap the gaze
origin has and worth avoiding twice: brows and lids move together constantly,
so a brow that jumped on every blink would read as a tic. Rest pose from the
take median rather than the neutral frame, for the reason gaze learned the hard
way - that frame is picked by minimum mouth aperture and says nothing about the
brows.

Two correspondences resolved from geometry, not declared: which ring is which
brow, and which end is the outer one. The second matters more - backwards, the
tilt mirrors and worry renders as its own opposite, which reads as a directed
performance choice rather than a bug and would never be questioned. Which EDGE
is upper is deliberately left unresolved: it traverses the same ring the other
way, an even-odd fill has no winding, and both ends still land on fixed slots.

Also fixes a bug from the exposure work: the live render applied exposure to
the plate and the mouth but not to the eyes, so on 2s the preview and the
export disagreed. A preview that disagrees with the export is the one bug this
tool cannot afford. perfIndex now exists as a named thing so the two paths
cannot drift apart again.

91 -> 105 assertions. Ground truth on all four synthetic brow poses, tilt
separating worry from anger by sign, a blink not faking a raise, and a shared
dwell never emitting a half-raised brow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:11:25 -04:00
Your Name
2e4406d579 Swap the test clip for the diner take
105 frames at 24fps from a screen recording, replacing IMG_8486. Ripped dense
on purpose now that exposure chooses the timing at render time.

Near-frontal this is not - it is a 3/4 view, so residual will run high and it
is the case the design doc says wants hand-drawn plates. That makes it the
better test clip for the eyes: there is real gaze movement in it, and the
out-of-plane bias on the projected iris offset is visible rather than
hypothetical.

frames/ is gitignored and regenerable: ./extract.sh <clip> 24

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:09 -04:00
Your Name
e497e9ac4d Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.

Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.

Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.

Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.

The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.

Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.

The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.

Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.

Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.

Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.

41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
Your Name
c2603102da readme additions 2026-09-24 15:49:05 -04:00
f827757f7d Become arthur: a standalone suite, not an Animator Pro front-end
The test renderer turned out to be the product. Everything that decides how the
work looks - stabilisation, reduction, timing, frame removal, palette - already
happens here, and the flat indexed output already reads the way it should.

The reason to leave is in the original design's own rule: never make a timing
decision that requires a full render to evaluate. Honouring that moved every
judgement out of Animator Pro, which left the host doing nothing but writing a
file, in exchange for modal UI, minutes-long renders, one-level undo, FLX delta
invariants, a single tween state and a cel singleton.

What does NOT change is the constraint. 320x200, indexed palette, flat fills,
no antialiasing - inherited, but load-bearing rather than accidental. The
rasteriser writes palette indices and expands to RGBA only at the end precisely
so nothing can soften an edge. Modern conveniences belong in the workflow.

Adds docs/design.md: the principles, carried over without the Poco/FLX/cel
machinery, plus architecture and an honest list of what is missing - the
largest gap being that plates still have nowhere to be drawn.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:47:41 -04:00
dc0a3f34e6 Make the mouth lead legible and assertable
"It feels like it does nothing" is indistinguishable from "it does nothing",
and at 24fps a lead of 1 is 42ms - small enough to reasonably doubt. So the
shift is now provable and visible rather than taken on trust:

- shiftIndex is a pure exported function with assertions covering identity,
  both directions, and clamping at each end
- the frame label always shows the mouth frame, not only when shifted, so the
  number can be watched diverging from f; non-zero leads also report in ms
- the slider readout carries an explicit + sign
- the stabilised pane draws the unshifted contour as a dark-green ghost when a
  lead is set, so the offset is something you can see

Also: leadIndex called opts() on every invocation - ~14 DOM reads, once per
strip thumbnail, so ~1000 per redraw at 74 frames. It reads a cached scalar now.

The "different pose" assertion checks the whole track rather than one pair:
synthetic poses hold for nine-frame beats, so a single pair can legitimately be
identical while the shift works correctly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:38:02 -04:00
f9b8ec8617 Fix teeth vertex slider driving the lip vertex count
opts() declared `verts` twice - once from the lip slider and again from the
teeth slider. Duplicate keys in an object literal are silent in JS and the last
one wins, so the teeth vertex control was quietly setting the lip vertex budget
while the lip control did nothing at all.

Renamed to teethVerts, with interior.js reading it under that name.

selftest now parses the opts() literal and fails on duplicate keys, verified to
catch this exact case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:36:04 -04:00
d18791e6ad Mouth lead: nudge performance tracks ahead of the clock
Smoothing does not literally delay anything - a centred moving average has zero
phase lag - but it blurs onsets, so the visually salient moment of a mouth
opening moves later even though the mean does not. Animators also draw mouth
shapes a frame or two ahead of the sound as standard practice, so this is the
normal control rather than a workaround.

Only the performance tracks shift; the head and audio stay put, since it is the
mouth that should anticipate. In photo-underlay modes the vector mouth will
therefore no longer match the frame behind it, which is expected.

The lead is baked into the exported take - key f carries the pose from source
frame f+lead - so the Animator Pro renderer never needs to know about it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:34:08 -04:00
a146d271b6 Cross-check el() ids against index.html in the selftest
A knob wired in app.js but missing from the markup throws during wiring, aborts
the module and leaves a blank page - a symptom pointing nowhere near its cause.
It has now happened twice, so it gets a check rather than vigilance: the
selftest fetches both files and compares the id sets in each direction.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:25:39 -04:00
9a11eeabc1 Teeth as an extracted blob contour, not a clipped band
The band filled the mouth because a band is the wrong reduction: the bright
region is a blob, and reading it as "everything above a line" throws the shape
away.

Extracting a contour reintroduces the vertex-correspondence problem that made
me avoid it, but for a blob there is a way out. Radial sampling from the
centroid along N fixed directions makes vertex k always mean "the extent in
direction k": correspondence holds by construction, the count is fixed, and
temporal smoothing cannot reorder anything. It also yields a star-shaped
reduction, which suits flat colour.

Tongue rejection, which the band had no way to express:
- pixels red relative to their own brightness are dropped (teeth are neutral)
- component choice is biased toward the top of the cavity, since area alone
  picks the tongue when the mouth is wide
- separate inner and outer controls: cavity erode pulls the sampled region off
  the lip edge, blob grow/erode resizes the found blob

Also: a knob wired in app.js but missing from index.html threw during wiring
and left a blank page with nothing useful in the console - which is exactly
what happened to teethDwell in the previous commit. el() now names the missing
id, and window.onerror surfaces it in the status line.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:24:52 -04:00
c6daf827d1 Fix teeth band filling the whole mouth
Three causes, all of them mine:

Otsu always returns a split, including on a homogeneous region - given a dark
cavity with no teeth it invents a threshold and calls half the pixels bright.
The gate is now the separation between the two class means, which is the only
thing that says whether the split means anything. Coverage was the wrong
signal: it is high both when the mouth is full of teeth and when the region is
uniformly dark and Otsu has split noise.

The row scan tracked the last qualifying row anywhere rather than where the run
from the top stops, so one bright row near the bottom - a lit lower lip inside
the ring - pushed the line to full height. It now breaks at the first failing
row once the run has started.

MediaPipe's inner lip landmarks sit slightly outside the real opening, so the
sampled region included lip pixels, which are bright and sit exactly at the
boundary where they do most damage. The ring is now eroded toward its centroid
before sampling, with the amount exposed as a knob.

Adds a diagnostic panel showing the sampled crop, pixels above threshold, and
the resolved line, because tuning this from numbers alone does not tell you
whether the region being measured is even the right region.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:14:49 -04:00
941022b69f Teeth from mouth-interior image content
MediaPipe has no landmarks inside the lips, so teeth have to come from pixels.
Tracing the bright blob would give a new contour every frame with no vertex
correspondence - the exact boil docs/roto-puppet.md warns about.

So the measurement yields a scalar, not a shape: Otsu within the cavity, scanned
from the top for where the bright run stops, giving one line height per frame.
The teeth polygon is the inner lip ring clipped to that line, so the silhouette
is always the mouth's own shape and cannot disagree with the lips around it,
and the only per-frame variable is a single number that smooths trivially.

Presence gets hysteresis and minimum dwell, as plate selection does: a teeth
block blinking on and off for single frames is worse than one simply absent.

Tongue is not implemented. The same scalar approach would apply, gated on
redness rather than brightness, but it is not visible in the test footage - the
cavity reads dark with a bright upper-teeth band and nothing else.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:12:09 -04:00
29b0bd6690 Fix horizontal stretch from MediaPipe's anisotropic normalised space
MediaPipe normalises x by image WIDTH and y by image HEIGHT, so for a 1080x1920
clip one unit of x is 1080px and one unit of y is 1920px. makeXform applied a
single scale to both, stretching everything horizontally by H/W - 1.78x on this
footage. The photo underlay looked equally squashed because frameAffine divided
x by imgW, matching the equally wrong vector shapes rather than disagreeing
with them.

Fixed at ingest: landmarks convert to an isotropic space whose unit is one image
height (x *= W/H), so equal numbers mean equal pixels everywhere downstream.
Pixel mapping follows - both axes divide by imgH.

This also silently fixes head roll. fitSimilarity was fitting a rotation in a
sheared space, so the "similarity" it recovered was not one, and stabilisation
of rolled heads was subtly wrong.

selftest: a shape circular in pixel space must stay circular in raster space,
checked at 1080x1920, 1920x1080 and 640x640. Fails at ratio 1.78 without the
conversion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:03:12 -04:00
de90bfd907 Registered photo underlay as the plate reference
The generated face oval was never going to be good enough to draw from:
MediaPipe's face oval is the FACE boundary, cut at the hairline and excluding
hair, ears, jaw and neck, so it is an egg by construction. Segmentation would
give a real head outline but costs a 16MB model and per-frame inference for a
shape that gets replaced by a drawing anyway.

So the plate layer becomes switchable, and the useful modes are photographic:
the source frame mapped into raster space through the same transform chain the
contours go through. Registration is the whole point - the head sits still and
a drawing traced from the underlay is already aligned to the mouth. An
unregistered underlay would be decoration.

- underlay.js: pixel->raster affine (a general affine, since MediaPipe
  normalises x by width and y by height), registered draw, palette posterise
- plate modes: photo / photo dim / posterized / oval / oval+photo / none, B cycles
- worksheet cells are registered composites rather than raw crops
- Save frame 4x writes a 1280x800 PNG to draw on
- selftest: FACE_OVAL simplicity, which was never asserted; a wrong ordering
  there reads as a lumpy plate rather than an obvious bowtie

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 14:57:59 -04:00
5ac919b9d6 Audio-clocked playback for judging sync
Lip sync cannot be judged silently. extract.sh now pulls the audio track and
writes manifest.json alongside the frames; the page reads the true extraction
rate from it rather than assuming one, since a guessed fps desynchronises
picture from sound - the one thing this view exists to show.

Audio is the clock: frame = floor(currentTime * fps). A slow render loop drops
frames instead of drifting, and half/quarter speed work via playbackRate with
the picture following for free. Scrubbing, stepping and clicking a thumbnail
all seek the audio too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 14:53:09 -04:00
a082208ad5 Frame removal for plates; mouth keeps every frame
Sparseness was being applied for two different reasons at once. Aesthetic
sparseness is set by the extraction rate; labour sparseness only binds on the
plate, because a human draws each one. The mouth is traced and therefore free,
and in limited animation lip sync is routinely the densest element - on 1s
while heads hold on 2s and 3s.

So: the mouth gets a key on every frame, and the frame strip is now the
editing surface for deciding which frames need their own plate drawing. All
frames start kept; delete the ones you don't want.

- strip of face-cropped thumbnails, keep/drop per frame, keyboard driven
- worksheet panel lists the drawings needed and the range each one holds
- Suggest runs error-tolerance decimation on head pose as a starting point
- export writes sparse plate keys + dense mouth keys, with a hold manifest
- smoothContours: bounded exception to "never smooth the contour", which held
  only while keys were sparse enough to reject detector noise by sampling
- averages are now a RADIUS in frames: 0 is off, 1 is +-1
- GPU delegate falls back to CPU instead of failing
- #synth / #frames autorun for headless smoke tests

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 14:51:15 -04:00
7a6bdea157 roto: video -> take file builder with interactive tuning
Analysis half of the pipeline in docs/roto-puppet.md. Stabilises a face out
of a clip via a similarity fit on rigid landmarks, reduces the lip contour to
a fixed vertex budget, selects sparse keys on velocity minima, and previews
the result as flat indexed fills so timing can be judged without an Animator
Pro render.

- landmarks.js  ordered lip/oval rings; slot position is vertex identity
- mathutil.js   closed-form 2D similarity, Procrustes mean, transform smoothing
- pipeline.js   stabilise -> subsample -> key-select
- raster.js     indexed scanline fill, no antialiasing
- take.js       take-file writer
- selftest.js   29 assertions, incl. ring simplicity at every vertex budget

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 14:38:07 -04:00