arthur/docs/design.md

316 lines
16 KiB
Markdown
Raw Normal View History

# arthur — design
A small animation suite for turning live-action video into work that reads as
hand-authored, in the idiom of *Another World* (Éric Chahi, 1991): flat shapes,
a tiny palette, hard edges, motion carried by silhouette.
## The aesthetic is a representation, not a filter
Chahi did not process video. He shot reference footage and hand-traced polygons
over it in a custom editor; the engine stored and replayed polygon lists, never
bitmaps. Three properties follow, and all three are load-bearing:
1. **Temporal identity.** The same shape, with the same vertex count and vertex
order, *edited* across frames. A shape re-detected independently each frame
produces a new contour every frame. That boils, and boiling reads as "filter"
within half a second no matter how good the individual shapes are.
2. **Stylised timing.** Poses held, and for most parts a hard cut between them
rather than an interpolation. A distinct pose on every frame of 24fps footage
looks like video even when every pose is a polygon.
3. **Authored colour.** A small fixed ramp with two or three tones per part,
chosen by a person. Sampling colour from the source produces a pixel-art
filter immediately and irrecoverably.
None of those are computer-vision problems. That is the whole argument about
where CV belongs.
## The constraint is the point
320×200, indexed palette, flat fills, no antialiasing. That is inherited from
Animator Pro, where this work started, but it is not an accident to be
modernised away — it is why the output looks right. The rasteriser writes palette
indices into a byte buffer and expands to RGBA only at the end, precisely so no
canvas antialiasing can soften an edge.
Modern conveniences belong in the *workflow* — instant feedback, audio, real
undo, scrubbing, layers. Not in the output.
## CV tracks; it does not draw
- **Its job.** Where the rigid features of the face are, what the head's
transform is, where the lip contour runs, which frame is a motion extreme.
- **Its non-job.** Deciding what the shapes are. A face mesh offers 468 points;
an animator's mouth is eight. **That decimation ratio is the style.** It is a
decision encoded in the tool, not a measurement extracted from footage.
Every number crossing from analysis into rendering is a *parameter* or a
*correspondence*. Every number determining how something looks comes from the
artist or a fixed authored table.
## Two kinds of part
| Kind | Source | Vocabulary | Interp |
| --- | --- | --- | --- |
| **Plate** — head, hair, body | Hand-drawn | Closed: a few drawings per character | hold |
Brows: traced ring, quantised raise A brow at 320x200 is fourteen pixels wide and three tall. Its shape carries almost nothing at that size; its height above the eye carries the expression, and a brow raise is the most legible beat on a face. So the ring is traced and the height is quantised - the split the eyes already got, where the lid is a traced feature and the iris a quantised primitive. The decomposition is the point. The traced ring already contains the real height, so adding a quantised raise on top would move the brow twice. The height is measured OUT of the ring, quantised, and put back, so the shape that renders is his at a height that snaps between a few levels and holds. Measured at both ends rather than as one number, because raise and tilt are different expressions out of one mechanism: both ends up is surprise, inner up alone is worry, inner down is anger. They share a dwell - the gaze quantiser, renamed quantizeSnap now that it has two callers - so the brow hits its pose in one frame instead of crawling into it with one end arriving before the other. Measured against the eye's corner midpoint, never its lid. Same trap the gaze origin has and worth avoiding twice: brows and lids move together constantly, so a brow that jumped on every blink would read as a tic. Rest pose from the take median rather than the neutral frame, for the reason gaze learned the hard way - that frame is picked by minimum mouth aperture and says nothing about the brows. Two correspondences resolved from geometry, not declared: which ring is which brow, and which end is the outer one. The second matters more - backwards, the tilt mirrors and worry renders as its own opposite, which reads as a directed performance choice rather than a bug and would never be questioned. Which EDGE is upper is deliberately left unresolved: it traverses the same ring the other way, an even-odd fill has no winding, and both ends still land on fixed slots. Also fixes a bug from the exposure work: the live render applied exposure to the plate and the mouth but not to the eyes, so on 2s the preview and the export disagreed. A preview that disagrees with the export is the one bug this tool cannot afford. perfIndex now exists as a named thing so the two paths cannot drift apart again. 91 -> 105 assertions. Ground truth on all four synthetic brow poses, tilt separating worry from anger by sign, a blink not faking a raise, and a shared dwell never emitting a half-raised brow. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:11:25 -04:00
| **Feature** — mouth, lids, brows | Rotoscoped from landmarks | Open: derived from this take | hold |
| **Interior** — mouth interior, teeth | Image content within a feature | Open | hold |
| **Primitive** — iris | Landmark centroid as a disc | Quantised | hold |
Brows: traced ring, quantised raise A brow at 320x200 is fourteen pixels wide and three tall. Its shape carries almost nothing at that size; its height above the eye carries the expression, and a brow raise is the most legible beat on a face. So the ring is traced and the height is quantised - the split the eyes already got, where the lid is a traced feature and the iris a quantised primitive. The decomposition is the point. The traced ring already contains the real height, so adding a quantised raise on top would move the brow twice. The height is measured OUT of the ring, quantised, and put back, so the shape that renders is his at a height that snaps between a few levels and holds. Measured at both ends rather than as one number, because raise and tilt are different expressions out of one mechanism: both ends up is surprise, inner up alone is worry, inner down is anger. They share a dwell - the gaze quantiser, renamed quantizeSnap now that it has two callers - so the brow hits its pose in one frame instead of crawling into it with one end arriving before the other. Measured against the eye's corner midpoint, never its lid. Same trap the gaze origin has and worth avoiding twice: brows and lids move together constantly, so a brow that jumped on every blink would read as a tic. Rest pose from the take median rather than the neutral frame, for the reason gaze learned the hard way - that frame is picked by minimum mouth aperture and says nothing about the brows. Two correspondences resolved from geometry, not declared: which ring is which brow, and which end is the outer one. The second matters more - backwards, the tilt mirrors and worry renders as its own opposite, which reads as a directed performance choice rather than a bug and would never be questioned. Which EDGE is upper is deliberately left unresolved: it traverses the same ring the other way, an even-odd fill has no winding, and both ends still land on fixed slots. Also fixes a bug from the exposure work: the live render applied exposure to the plate and the mouth but not to the eyes, so on 2s the preview and the export disagreed. A preview that disagrees with the export is the one bug this tool cannot afford. perfIndex now exists as a named thing so the two paths cannot drift apart again. 91 -> 105 assertions. Ground truth on all four synthetic brow poses, tilt separating worry from anger by sign, a blink not faking a raise, and a shared dwell never emitting a half-raised brow. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:11:25 -04:00
| **Scalar** — brow raise, gaze | One number out of a feature | Quantised | hold |
The last row took the longest to see. A brow is a feature *and* a scalar: the
ring is traced because the shape should be his, but at three pixels tall the
shape carries almost nothing while the height above the eye carries the
expression. So the height is measured out of the traced ring, quantised, and put
back. Extracting the scalar without removing it first would move the part twice,
because the traced ring already contains the height.
The asymmetry is deliberate, and it is the opposite choice in each case.
A **closed vocabulary is wrong for the mouth.** Pre-authored mouth shapes are
never *this performance's* shapes, and that genericness is what the project
exists to avoid. So the mouth gets an open vocabulary derived from footage, and
the stylisation applies to its **timing** instead.
A **closed vocabulary is right for the head.** The head carries structure, not
performance. A handful of hand-drawn angles is better, because you drew them —
and it means the head never needs segmentation, contour tracking, or boil
avoidance.
## Two kinds of sparseness
Eyes: lids, blinking, line of sight Three parts per eye, stacked the way the mouth is - dark lash ring, sclera inside it, iris inside that, square pupil in the iris. A blink then costs nothing: when the lid shuts the traced ring goes flat and the lash line collapses to a lens, which is a closed eye, drawn correctly, for free. Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a quantised position - and that is where the stylisation lives. Line of sight. Gaze is the iris centre relative to the midpoint of the eye's two corners, in units of corner distance. Both corners are in RIGID, so the origin and the scale are immune to the performance being measured; against the lid ring's centroid instead, every blink would drag the origin down and fake a glance at the floor on exactly the frames where the eye is most visible. Both eyes share one gaze - at this size the difference between the two measurements is noise, not vergence, and independent per-eye noise reads as wall-eyed immediately. Openness stays per-eye so a wink survives. Gaze is then quantised to a pixel grid with a dwell, which is not a stylisation imposed on the truth: real eyes move in saccades, and the smooth drift left in the measurement is tracker noise plus head-compensation error. Snapping to a grid removes the noise and recovers the saccade in one operation. The iris is placed in the frame of the already-smoothed, already-subsampled lid ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any even budget keeps them at 0 and n/2 - so it cannot drift relative to its own eye. Size is authored from the take mean, never remeasured per frame: a radius that breathes by a fraction of a pixel flickers a pixel on and off around the whole silhouette. iris anchor toggles steady/free/locked, because how much the eye wanders turns out to be an aesthetic choice and not only a correctness one. Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not have: blink hold. A blink is one frame at 12fps and a single frame of closed eye reads as a dropped frame, so once the eye shuts it stays shut long enough to be legible. Detection accuracy is not the problem; legibility is. The pupil is a square because at three pixels a circle is a plus sign with the corners gnawed off, and it changes shape as it moves. Drawn from a rounded centre shared with the iris so it is exactly its nominal size on every frame. Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro would: the lid crops the iris at extreme gaze for free, so nothing has to clamp the gaze, which would flatten the performance at the extremes that carry it. Which iris block belongs to which eye is RESOLVED from geometry, not declared. A swap looks almost right - each eye still has a disc roughly where it belongs - so it survives an eyeball and then reads as a subtly wall-eyed character forever. Voted across every frame; the test feeds a deliberately swapped track. Also: exposure. Aesthetic sparseness was set by the extraction rate, which made the timing a property of a directory of PNGs - auditioning 12 against 24 meant re-ripping and re-detecting the whole clip. It is now a render-time grid, on 1s/2s/3s/4s, so the dense track keeps everything and the audio clock is untouched. The take format already carried an exposure field; it was never driven. Everything rides the same grid, because a head cutting on the odd frames while the mouth cuts on the even ones reads as two performances laid over each other. 41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a swapped track, a blink does not fake a change of gaze, a stencilled disc cannot spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure never reads a pose from the future. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
Conflating these was the original design error. **Aesthetic** sparseness is the
rate the picture changes at. **Labour** sparseness is a human drawing each one,
and it binds only on the plate.
Aesthetic sparseness used to be set by the extraction rate — rip at 12 and the
timing is chosen. That was wrong in a small way: it makes the timing a property
of a directory of PNGs, so auditioning 12 against 24 means re-ripping the clip
and re-running detection over the whole of it, and the decision you most want to
play with is the one that costs the most to change. Rip at the camera's rate and
quantise at render time instead — an **exposure** grid, on 1s, 2s, 3s — so the
dense track keeps everything, the audio clock is untouched, and the timing is a
dropdown rather than a re-rip. The take format already carried an `exposure`
field for this; it was simply never driven.
So the mouth keeps **every** frame — it is traced, and therefore free. In limited
animation lip sync is routinely the densest element, on 1s, while heads hold on
2s and 3s.
The other half is frame removal: which frames need their own plate drawing.
Everything starts kept; delete what you don't want. Automatic suggestion runs
error-tolerance decimation over head pose as a starting point, then you
hand-correct.
## Stabilisation
A rotoscoped mouth is only reusable if expressed independently of where the head
was. That inverts the obvious composition:
```text
mouth_local(t) = anchor(t)⁻¹ · mouth_world(t)
```
- **Rigid landmarks only** for the fit: eye corners, nose bridge, nose tip.
Including a feature that moves bleeds performance into the stabilisation.
- **Similarity, not affine or homography.** The extra degrees of freedom absorb
head rotation as distortion and smear it into the mouth. Four DOF removes
exactly translation, roll and depth scale, leaving yaw and pitch as a
measurable residual.
- **Reference is the mean** configuration over the shot, not frame zero.
- **Smooth the transform, never the contour** — with one bounded exception,
below.
- Landmarks convert to an **isotropic** space first (unit = one image height).
MediaPipe normalises x by width and y by height, so its space is stretched; a
"similarity" fitted there is not one.
Out-of-plane rotation cannot be removed by any 2D transform. The answer is not a
better transform: it is to draw the head at each angle and let each drawing
declare where its mouth sits. Foreshortening becomes authored metadata.
## Hold versus interpolate
Another World had no inbetweening engine: polygon sets played back frame by
frame at a low rate. For parts that should read as hand-animated snaps — mouths
above all — **cut between keys, do not interpolate.** A tweened mouth is rubbery
and reads as puppet software immediately.
**Fixed topology is non-negotiable for cut parts.** Vertex *meanings* must be
stable across every key, so a closed mouth and a wide-open mouth are the same
polygon at different positions. Interpolation partially hides ordering drift;
cutting exposes it completely, as static.
## Deriving shapes from pixels without boiling
Some things have no landmarks — teeth, tongue, anything inside the lips. The
hazard is vertex correspondence: a traced contour reorders between frames.
Two escapes, both used:
- **Extract a scalar, not a shape**, where the shape can be derived from
geometry you already trust.
- **Radial sampling** where a real contour is needed. March outward from the
blob's centroid along N fixed directions; vertex *k* is then always "the extent
in direction *k*". Correspondence holds by construction, the count is fixed,
temporal smoothing is well defined, and the star-shaped result suits flat
colour.
Eyes: lids, blinking, line of sight Three parts per eye, stacked the way the mouth is - dark lash ring, sclera inside it, iris inside that, square pupil in the iris. A blink then costs nothing: when the lid shuts the traced ring goes flat and the lash line collapses to a lens, which is a closed eye, drawn correctly, for free. Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a quantised position - and that is where the stylisation lives. Line of sight. Gaze is the iris centre relative to the midpoint of the eye's two corners, in units of corner distance. Both corners are in RIGID, so the origin and the scale are immune to the performance being measured; against the lid ring's centroid instead, every blink would drag the origin down and fake a glance at the floor on exactly the frames where the eye is most visible. Both eyes share one gaze - at this size the difference between the two measurements is noise, not vergence, and independent per-eye noise reads as wall-eyed immediately. Openness stays per-eye so a wink survives. Gaze is then quantised to a pixel grid with a dwell, which is not a stylisation imposed on the truth: real eyes move in saccades, and the smooth drift left in the measurement is tracker noise plus head-compensation error. Snapping to a grid removes the noise and recovers the saccade in one operation. The iris is placed in the frame of the already-smoothed, already-subsampled lid ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any even budget keeps them at 0 and n/2 - so it cannot drift relative to its own eye. Size is authored from the take mean, never remeasured per frame: a radius that breathes by a fraction of a pixel flickers a pixel on and off around the whole silhouette. iris anchor toggles steady/free/locked, because how much the eye wanders turns out to be an aesthetic choice and not only a correctness one. Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not have: blink hold. A blink is one frame at 12fps and a single frame of closed eye reads as a dropped frame, so once the eye shuts it stays shut long enough to be legible. Detection accuracy is not the problem; legibility is. The pupil is a square because at three pixels a circle is a plus sign with the corners gnawed off, and it changes shape as it moves. Drawn from a rounded centre shared with the iris so it is exactly its nominal size on every frame. Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro would: the lid crops the iris at extreme gaze for free, so nothing has to clamp the gaze, which would flatten the performance at the extremes that carry it. Which iris block belongs to which eye is RESOLVED from geometry, not declared. A swap looks almost right - each eye still has a disc roughly where it belongs - so it survives an eyeball and then reads as a subtly wall-eyed character forever. Voted across every frame; the test feeds a deliberately swapped track. Also: exposure. Aesthetic sparseness was set by the extraction rate, which made the timing a property of a directory of PNGs - auditioning 12 against 24 meant re-ripping and re-detecting the whole clip. It is now a render-time grid, on 1s/2s/3s/4s, so the dense track keeps everything and the audio clock is untouched. The take format already carried an exposure field; it was never driven. Everything rides the same grid, because a head cutting on the odd frames while the mouth cuts on the even ones reads as two performances laid over each other. 41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a swapped track, a blink does not fake a change of gaze, a stencilled disc cannot spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure never reads a pose from the future. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
## Every part is measured in the frame of the thing it is attached to
The mouth is expressed against the head. The iris is expressed against its own
eye — specifically against the midpoint of that eye's two corners, in units of
corner distance. Both corners are rigid landmarks, so the origin and the scale
of the measurement are immune to the performance being measured. Against the lid
ring's centroid instead, every blink would drag the origin down and fake a glance
at the floor on exactly the frames where the eye is most visible.
The rule generalises: **measure a feature in a frame built only from landmarks
that do not move with it.** It is the same argument as "rigid landmarks only" for
the anchor fit, one level down.
There is a tempting over-application. An eye can be pinned into a fixed socket
fitted to its corners' mean over the shot, which removes the residual wobble a
2D similarity cannot — and it is wrong. That residual is real motion of the eye
relative to the head, it is still there in the footage, and removing it leaves
the drawn eyes hanging still over a registered photo whose eyes are moving. A
part must track the face in the same space the underlay is drawn in. The wobble
is a job for the bounded contour average below, not for a second anchor.
Placement follows from the same idea. The iris is drawn in the frame of the
already-smoothed, already-subsampled lid ring, read off the ring's own corner
vertices, so it cannot drift relative to the eye it sits inside and it inherits
the contour average for free. Size, by contrast, is authored from the take's
mean, never remeasured per frame: a radius that breathes by a fraction of a pixel
flickers a pixel on and off around the whole silhouette.
## The indexed buffer is its own stencil
Parts that nest — iris inside sclera, pupil inside iris — clip by colour key:
paint only where the buffer already holds the parent's index. This is how
Animator Pro would do it, it costs one comparison per pixel, and it composes
transitively, so a blink takes the right bite out of the pupil without anything
computing where.
It also removes a temptation. Without a stencil the gaze has to be clamped to
keep the iris inside the lid, and a clamp flattens the performance at exactly the
extremes that carry it.
## Quantisation can be the truthful choice
Gaze snapped to a pixel grid with a dwell is the "Primitive — quantised" row of
the part table, and it looks like a stylisation imposed on a continuous
measurement. It is not. Real eyes move in saccades: hold a fixation, jump, hold.
The smooth drift left in the measured signal is tracker noise plus
head-compensation error. Snapping to a grid and requiring a dwell removes the
noise and recovers the saccade in the same operation — the rare case where the
aesthetic rule and the physiology agree.
The count of distinct cells the iris ever occupies is the number the knobs exist
to control. Three or four is a character who looks at things; forty is an
unquantised iris sliding around.
## Some thresholds need a minimum duration, not just a dwell
A dwell delays a change until it has persisted, which is the right guard against
chatter and is what the teeth use. A blink needs the opposite guard as well. It
lasts 100–150ms — one frame at 12fps — and a single frame of closed eye reads as
a dropped frame rather than as a blink. Animators draw a blink over two or three
drawings for that reason, so once the eye shuts it must stay shut for a minimum
number of frames. Detection accuracy is not the problem; legibility is.
## Resolve correspondences from data when a wrong guess is survivable
The refined mesh appends ten iris points, five per eye, and the upstream
left/right naming is viewer-relative in some documentation and subject-relative
in others. Swapping them looks *almost* right — each eye still has a disc roughly
where it belongs — so the error survives inspection and then reads as a subtly
wall-eyed character for the life of the project.
A hardcoded table is the wrong shape for a fact like that. Voting each block's
distance to each eye's corner midpoint across every frame settles it from the
geometry, cannot be got wrong, and keeps working if the model is renumbered. The
test feeds it a track built the other way round, because a resolver checked only
against the convention it was written for is checking nothing.
## The bounded smoothing exception
*Smooth the transform, never the contour* held while keys were sparse: sampling
at velocity minima rejected per-frame detector noise for free. With a key on
every frame it does not, so a short contour average is justified — the window
must stay **shorter than the shortest articulation worth keeping**. At 12fps,
mouth movement spans 3–6 frames and detector noise is per-frame, so ±1 separates
them and ±3 starts eating speech. At 24fps, double it.
## Performance tracks may lead the clock
A centred moving average has no phase lag, but it blurs onsets, so the visually
salient moment of a mouth opening moves later even though the mean does not.
Animators also draw mouth shapes a frame or two ahead of the sound as standard
practice. Only performance tracks shift; the head stays with the audio, because
it is the mouth that should anticipate.
## Colour discipline
- A small ramp. Another World ran 16 colours at 320×200.
- Two or three tones per part: base, shadow, occasionally a rim.
- No dithering, no antialiasing anywhere.
- Never sample colour from the source. Analysis may report *which* tone a region
should be; it must never report an RGB value.
- Plate art and generated parts share one palette, authored together.
## Three artifacts, not two
| Artifact | Cost | Regenerated when |
| --- | --- | --- |
| Dense track — landmarks and fitted anchor per frame | Minutes, once per shot | Re-shoot, or a model change |
| Take — parts, keys, kept frames | Instant | Every knob change |
| Render | — | Continuous |
Keep the dense track: re-keying is then a regenerate, never a re-trace. And
because the take is disposable, **hand corrections must not live in it** — they
belong in an override layer keyed by `(part, frame)`, applied on top at render
time.
## Architecture
The organising idea is an **indexed raster core**, a **part/track model**, and
**pluggable sources** feeding it. Sources so far are rotoscope and image-derived
interior; hand-drawn and procedural are the obvious next ones. Everything else —
stabilisation, key selection, frame removal, palette — operates on the track
model and is source-agnostic.
Current modules:
| Module | Role |
| --- | --- |
| `landmarks.js` | Index tables. Ring arrays are ordered traversals: slot position *is* vertex identity. |
Eyes: lids, blinking, line of sight Three parts per eye, stacked the way the mouth is - dark lash ring, sclera inside it, iris inside that, square pupil in the iris. A blink then costs nothing: when the lid shuts the traced ring goes flat and the lash line collapses to a lens, which is a closed eye, drawn correctly, for free. Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a quantised position - and that is where the stylisation lives. Line of sight. Gaze is the iris centre relative to the midpoint of the eye's two corners, in units of corner distance. Both corners are in RIGID, so the origin and the scale are immune to the performance being measured; against the lid ring's centroid instead, every blink would drag the origin down and fake a glance at the floor on exactly the frames where the eye is most visible. Both eyes share one gaze - at this size the difference between the two measurements is noise, not vergence, and independent per-eye noise reads as wall-eyed immediately. Openness stays per-eye so a wink survives. Gaze is then quantised to a pixel grid with a dwell, which is not a stylisation imposed on the truth: real eyes move in saccades, and the smooth drift left in the measurement is tracker noise plus head-compensation error. Snapping to a grid removes the noise and recovers the saccade in one operation. The iris is placed in the frame of the already-smoothed, already-subsampled lid ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any even budget keeps them at 0 and n/2 - so it cannot drift relative to its own eye. Size is authored from the take mean, never remeasured per frame: a radius that breathes by a fraction of a pixel flickers a pixel on and off around the whole silhouette. iris anchor toggles steady/free/locked, because how much the eye wanders turns out to be an aesthetic choice and not only a correctness one. Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not have: blink hold. A blink is one frame at 12fps and a single frame of closed eye reads as a dropped frame, so once the eye shuts it stays shut long enough to be legible. Detection accuracy is not the problem; legibility is. The pupil is a square because at three pixels a circle is a plus sign with the corners gnawed off, and it changes shape as it moves. Drawn from a rounded centre shared with the iris so it is exactly its nominal size on every frame. Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro would: the lid crops the iris at extreme gaze for free, so nothing has to clamp the gaze, which would flatten the performance at the extremes that carry it. Which iris block belongs to which eye is RESOLVED from geometry, not declared. A swap looks almost right - each eye still has a disc roughly where it belongs - so it survives an eyeball and then reads as a subtly wall-eyed character forever. Voted across every frame; the test feeds a deliberately swapped track. Also: exposure. Aesthetic sparseness was set by the extraction rate, which made the timing a property of a directory of PNGs - auditioning 12 against 24 meant re-ripping and re-detecting the whole clip. It is now a render-time grid, on 1s/2s/3s/4s, so the dense track keeps everything and the audio clock is untouched. The take format already carried an exposure field; it was never driven. Everything rides the same grid, because a head cutting on the odd frames while the mouth cuts on the even ones reads as two performances laid over each other. 41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a swapped track, a blink does not fake a change of gaze, a stencilled disc cannot spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure never reads a pose from the future. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
| `pipeline.js` | …also eye openness, gaze, blink resolution and the iris pairing vote. |
| `mathutil.js` | Similarity fit, Procrustes mean, temporal smoothing. |
| `pipeline.js` | Stabilise → subsample → key-select → frame-removal. |
| `interior.js` | Teeth from image content: Otsu, morphology, components, radial contour. |
| `underlay.js` | Registered photo reference and palette posterisation. |
| `raster.js` | Indexed scanline fill. No antialiasing, by construction. |
| `take.js` | Take-file writer. |
| `synth.js` | Synthetic landmarks, so everything below detection is testable without a video. |
| `selftest.js` | Assertions, including source-level wiring checks. |
## Not built yet
- **A paint surface.** The plates have nowhere to be drawn. This is the largest
gap between "tool" and "suite": a pixel paint canvas with onion skin, palette
constraint, and the registered underlay behind it.
Brows: traced ring, quantised raise A brow at 320x200 is fourteen pixels wide and three tall. Its shape carries almost nothing at that size; its height above the eye carries the expression, and a brow raise is the most legible beat on a face. So the ring is traced and the height is quantised - the split the eyes already got, where the lid is a traced feature and the iris a quantised primitive. The decomposition is the point. The traced ring already contains the real height, so adding a quantised raise on top would move the brow twice. The height is measured OUT of the ring, quantised, and put back, so the shape that renders is his at a height that snaps between a few levels and holds. Measured at both ends rather than as one number, because raise and tilt are different expressions out of one mechanism: both ends up is surprise, inner up alone is worry, inner down is anger. They share a dwell - the gaze quantiser, renamed quantizeSnap now that it has two callers - so the brow hits its pose in one frame instead of crawling into it with one end arriving before the other. Measured against the eye's corner midpoint, never its lid. Same trap the gaze origin has and worth avoiding twice: brows and lids move together constantly, so a brow that jumped on every blink would read as a tic. Rest pose from the take median rather than the neutral frame, for the reason gaze learned the hard way - that frame is picked by minimum mouth aperture and says nothing about the brows. Two correspondences resolved from geometry, not declared: which ring is which brow, and which end is the outer one. The second matters more - backwards, the tilt mirrors and worry renders as its own opposite, which reads as a directed performance choice rather than a bug and would never be questioned. Which EDGE is upper is deliberately left unresolved: it traverses the same ring the other way, an even-odd fill has no winding, and both ends still land on fixed slots. Also fixes a bug from the exposure work: the live render applied exposure to the plate and the mouth but not to the eyes, so on 2s the preview and the export disagreed. A preview that disagrees with the export is the one bug this tool cannot afford. perfIndex now exists as a named thing so the two paths cannot drift apart again. 91 -> 105 assertions. Ground truth on all four synthetic brow poses, tilt separating worry from anger by sign, a blink not faking a raise, and a shared dwell never emitting a half-raised brow. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:11:25 -04:00
- A tongue.
- Plate libraries with per-plate mouth slots.
- Real performer→character calibration (currently identity).
- The override layer.
- Export beyond the take file — FLI/FLC would connect back to the lineage and is
a simple format.
## Lineage
This began as an analysis front-end for Animator Pro, whose render path was the
intended destination. `Animator-Pro-SDL/docs/roto-puppet.md` documents that
design and what the Poco/FLX machinery can and cannot do. The reason for leaving
is in that document's own rule — *never make a timing decision that requires a
full render to evaluate* — which, once honoured, left the host with nothing to do
but write the file.