A sketch, and labelled as one. It exists to test whether the aesthetic holds when a human draws the background rather than the tracker deriving it, and it is meant to be replaced by a real paint surface with onion skin and undo. Kept to one dependency-free module so throwing it away is a delete, not surgery. Cels are drawn on the frames that get their own drawing and hold until the next one - the same rule the plate follows, and literally the same lookup, so the two cannot disagree about what is on screen. Editing always targets the cel you can see, so you can scrub anywhere and keep drawing. Pen places vertices and closes on the first one. Edit drags vertices or whole shapes, shift-click inserts, alt-click removes. Layers stack front-at-top with per-layer colour, show/hide and reorder. Drag a strip thumbnail onto the canvas to seed this cel from that one, every layer, as a deep copy - sharing the point arrays would make two cels silently edit each other. Two rules are enforced rather than left to discipline: Colours are PALETTE INDICES, never RGB. Sampling colour from the source is the one move docs/design.md calls irrecoverable, and a paint tool is exactly where that discipline would leak, so the picker cannot express a colour outside the ramp. Vertices snap to the 320x200 grid. On a hard-edged indexed rasteriser a shape nudged by 0.4px moves an edge by a whole pixel or not at all depending on where it happens to land, so sub-pixel vertices shimmer instead of holding still. Drawings autosave to localStorage per take name. They are the only thing here a person made by hand; everything else regenerates. Not in the .take export yet. Also adds serve.py, a no-store dev server. python3 -m http.server sends Last-Modified and browsers cache ES modules on it hard enough that a reload serves a stale app.js against a fresh index.html: the new knobs appear, nothing wires them, no error fires, and it reads as "your feature does not work". That cost real time this session. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
370 lines
19 KiB
Markdown
370 lines
19 KiB
Markdown
# arthur
|
||
|
||
An animation suite for turning live-action video into work that reads as
|
||
hand-authored — flat shapes, a tiny palette, hard edges, motion carried by
|
||
silhouette, in the idiom of *Another World*.
|
||
|
||
It stabilises a face out of a clip, reduces the lip contour to a handful of
|
||
vertices, derives teeth from image content, lets you decide which frames need
|
||
their own hand-drawn head, and renders the result as flat indexed fills at
|
||
320×200 with no antialiasing.
|
||
|
||
The constraint is deliberate. 320×200 and an indexed palette are inherited from
|
||
Animator Pro, where this started, but they are why the output looks right —
|
||
modern conveniences belong in the workflow, not the output. See
|
||
[docs/design.md](docs/design.md).
|
||
|
||
## Run
|
||
|
||
```sh
|
||
python3 serve.py # from this directory, then open 127.0.0.1:8777
|
||
```
|
||
|
||
Use `serve.py`, not `python3 -m http.server`. The latter sends `Last-Modified`
|
||
and browsers cache ES modules on it hard enough that a reload serves a stale
|
||
`js/app.js` against a fresh `index.html` — new knobs appear in the markup, nothing
|
||
wires them, no error is raised, and the symptom reads as "the feature does not
|
||
work". `serve.py` is the same server with `no-store`.
|
||
|
||
Static files and ES modules — no build step, no dependencies beyond MediaPipe's
|
||
wasm, which is fetched from a CDN on first use.
|
||
|
||
**Synthetic take** needs no video and exercises everything below detection.
|
||
|
||
For real footage:
|
||
|
||
```sh
|
||
./extract.sh /path/to/clip.mov 12 # -> frames/*.png, audio.wav, manifest.json
|
||
```
|
||
|
||
then **Load frames**. MediaPipe's wasm is fetched from jsdelivr on first use;
|
||
`face_landmarker.task` is local.
|
||
|
||
Frames are pre-extracted rather than decoded in the page because browser video
|
||
seeking is approximate and `requestVideoFrameCallback` only delivers frames at
|
||
playback speed — neither gives a deterministic per-frame pass.
|
||
|
||
`manifest.json` records the true extraction rate. The page reads it rather than
|
||
assuming, because a guessed fps desynchronises audio from picture — and sync is
|
||
the one thing this view exists to show.
|
||
|
||
**Exposure** decides how often the picture gets a new drawing: rip at 24 and
|
||
render `on 2s` for 12, `on 3s` for 8. The dense track and the audio are
|
||
untouched, so it is a dropdown rather than a re-rip, and the export emits keys
|
||
only on the grid instead of the same pose twice. Everything rides the same grid
|
||
— mouth, eyes, teeth, plate — because a head cutting on the odd frames while the
|
||
mouth cuts on the even ones reads as two performances laid over each other.
|
||
|
||
**Audio is the playback clock**: `frame = floor(audio.currentTime * fps)`. A slow
|
||
render loop therefore drops frames instead of drifting, and ½x / ¼x work by
|
||
setting `playbackRate` with the picture following for free.
|
||
|
||
## Shooting for it
|
||
|
||
Near-frontal, good light, consistent scale, head reasonably still. Hold a neutral
|
||
closed mouth for a second at the top of the take: that frame is picked
|
||
automatically as the neutral and drives calibration and the placeholder plate.
|
||
|
||
Out-of-plane head rotation cannot be stabilised away by a 2D similarity
|
||
transform — the `residual` readout rises when it happens. The pipeline's answer
|
||
is hand-drawn head plates, which this tool does not yet do.
|
||
|
||
## Knobs
|
||
|
||
| Knob | What it does |
|
||
| --- | --- |
|
||
| vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. |
|
||
| exposure | How often the picture changes: on 1s, 2s, 3s, 4s. Rip dense, choose timing here. |
|
||
| mouth lead ±f | Shifts the performance tracks earlier against the audio and the head. `[` `]`. |
|
||
| contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. |
|
||
| anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform*. |
|
||
| closed-mouth cut | Aperture below which the mouth interior is emitted as `hidden`. |
|
||
| suggest tolerance | Max head movement before a new plate drawing is required. Affects **Suggest** only. |
|
||
|
||
## Teeth
|
||
|
||
MediaPipe has no landmarks inside the lips: the inner ring bounds the cavity and
|
||
everything within it is just pixels. So teeth come from the image.
|
||
|
||
The hazard is vertex correspondence — a traced contour reorders between frames
|
||
and boils. The way out for a blob specifically is **radial sampling**: march
|
||
outward from the blob's centroid along N fixed directions and take the last pixel
|
||
inside. Vertex *k* is then always "the extent in direction *k*", so correspondence
|
||
holds by construction, the vertex count is fixed, and temporal smoothing is well
|
||
defined with no reordering possible. It also produces a star-shaped reduction,
|
||
which is what flat blocks of colour want.
|
||
|
||
The **teeth measurement** panel shows exactly what is sampled: region dimmed,
|
||
kept pixels green, extracted contour amber. Tune against that, not the numbers.
|
||
|
||
| Knob | What it does |
|
||
| --- | --- |
|
||
| teeth contrast | Gate on the separation between the cavity's dark and bright class means. Otsu always returns *some* threshold, so this is what stops it inventing teeth in a dark mouth. |
|
||
| cavity erode | Pulls the sampled region in from the lip edge — MediaPipe's inner lip landmarks sit slightly outside the real opening, and lips are bright. |
|
||
| blob grow/erode | Resizes the found blob. An open pass always runs first to despeckle. |
|
||
| tongue reject | Drops pixels red relative to their own brightness. Teeth are near-neutral; tongue is not. |
|
||
| prefer upper | Biases component choice toward the top of the cavity. Area alone picks the tongue when the mouth is wide. |
|
||
| teeth vertices | Radial sample count. |
|
||
| teeth avg ±f | Temporal average over the contour. |
|
||
| teeth dwell | Frames a presence change must persist before it takes effect. |
|
||
|
||
**Tongue** as its own part would work the same way, gated on redness instead of
|
||
brightness and biased low rather than high. Not implemented: it is not visible in
|
||
the test footage, which reads as a dark cavity with a bright upper-teeth band.
|
||
|
||
## Eyes
|
||
|
||
Three parts per eye, stacked the way the mouth is: a dark **lash ring**, the
|
||
**sclera** inside it, and the **iris** inside that, with a square **pupil** in
|
||
the iris. The dark ring outside a pale interior is what makes a flat shape read
|
||
as an opening rather than a blob, and it is why a blink costs nothing — when the
|
||
lid shuts, the traced ring goes flat and the lash line collapses to a lens,
|
||
which is a closed eye, drawn correctly, for free.
|
||
|
||
The **lids are a feature**, rotoscoped like the mouth: head-local, a key on every
|
||
frame, the same `contour avg` knob. They track the face, because the face is what
|
||
they are attached to.
|
||
|
||
The **iris is a primitive** — a disc at a quantised position — and that is where
|
||
the stylisation is.
|
||
|
||
### Line of sight
|
||
|
||
Gaze is the iris centre relative to the **midpoint of the eye's two corners**,
|
||
in units of corner distance. Both corners are in `RIGID`, which is the point:
|
||
the origin and the scale are built only from landmarks that do not move under
|
||
performance. Measure against the lid ring's centroid instead and every blink
|
||
drags that centroid down and fakes a glance at the floor, on exactly the frames
|
||
where the eye is most conspicuous.
|
||
|
||
**Both eyes share one gaze.** At 320×200 an iris is a handful of pixels and its
|
||
centre comes from five landmarks on an eye twenty pixels wide, so the difference
|
||
between the two measurements is noise, not vergence — and independent per-eye
|
||
noise reads as wall-eyed immediately, which is the most expensive artefact on a
|
||
face. Openness stays per-eye, so a wink survives.
|
||
|
||
Then the gaze is **quantised to a pixel grid with a dwell**, which is not a
|
||
stylisation imposed on the truth: real eyes move in saccades, holding a fixation
|
||
and then jumping. The smooth drift left in the measurement is tracker noise plus
|
||
head-compensation error, so snapping to a grid and requiring a dwell removes the
|
||
noise and recovers the saccade in one operation. The readout reports how many
|
||
distinct cells the iris ever occupies — three or four is a character who looks at
|
||
things, forty is an unquantised iris sliding around.
|
||
|
||
The iris is drawn at the socket read back off the **already-smoothed, already-
|
||
subsampled lid ring** — slots 0 and 8 of a 16-slot ring are the corners, and
|
||
subsampling to any even budget keeps them at output indices 0 and `n/2`. So the
|
||
iris is placed in the frame of the exact polygon it sits inside and cannot drift
|
||
relative to its own eye. Size is authored from the take's mean eye width, not
|
||
remeasured per frame: a radius that breathes by a fraction of a pixel flickers a
|
||
pixel on and off around the whole silhouette.
|
||
|
||
The iris is **stencilled to the sclera** and the pupil to the iris — the indexed
|
||
buffer is its own clip mask, the way Animator Pro would do it. So the lid crops
|
||
the iris at extreme gaze automatically, and nothing needs to clamp the gaze,
|
||
which would flatten the performance at exactly the extremes that carry it.
|
||
|
||
### Blinking
|
||
|
||
Openness is the lid gap over the corner distance — normalised, so one threshold
|
||
carries across takes and faces. It gets hysteresis and a dwell like the teeth,
|
||
plus one knob the teeth do not have: **blink hold**. A blink is 100–150ms, which
|
||
is one frame at 12fps, and a single frame of closed eye reads as a dropped frame
|
||
rather than as a blink. Animators draw a blink over two or three drawings for
|
||
that reason, so once the eye shuts it stays shut for `hold` frames.
|
||
|
||
### The pupil is a square
|
||
|
||
At this size a pupil is three pixels across, and a circle of radius 1.5 is not a
|
||
circle — it is a plus sign with the corners gnawed off, and it changes shape as
|
||
it moves. A square that size is a deliberate mark that stays the same mark
|
||
wherever it lands. It is drawn from a rounded centre shared with the iris, so it
|
||
is exactly its nominal size on every frame instead of spilling to the next pixel
|
||
on some and not others.
|
||
|
||
### Which iris is which
|
||
|
||
The refined mesh appends ten iris points, five per eye, and MediaPipe's own
|
||
left/right naming is viewer-relative in some places and subject-relative in
|
||
others. Getting it backwards swaps the irises, which looks *almost* right — each
|
||
eye still has a disc roughly where it belongs — so it survives an eyeball and
|
||
then reads as a subtly wall-eyed character forever. The pairing is therefore
|
||
**resolved from the geometry**, by voting each block's distance to each eye's
|
||
corner midpoint across every frame, and the selftest feeds it a track built the
|
||
other way round to prove it actually looks.
|
||
|
||
| Knob | What it does |
|
||
| --- | --- |
|
||
| eye vertices | Lid ring vertex budget, off a 16-slot ring. |
|
||
| lash line | How far the dark ring sits outside the lid, in pixels. |
|
||
| blink cut | Openness below which the eye is shut. Normalised by corner distance. |
|
||
| blink hold | Minimum frames a blink stays on screen. A one-frame blink is a dropout. |
|
||
| blink dwell | Frames a change must persist. Usually 0 — unlike the teeth, a real blink *is* one frame. |
|
||
| gaze gain | Exaggerates or damps the throw. Measured excursion is small; a character usually wants more. |
|
||
| gaze step | The pixel grid the iris snaps to. 0 = off, and then dwell does nothing either. |
|
||
| gaze dwell | How long a new cell must hold before it takes. Together with step, this is what makes saccades. |
|
||
| iris size | Diameter as a percentage of eye width. |
|
||
| pupil | Square pupil in whole pixels. 0 = off. |
|
||
|
||
## Brows
|
||
|
||
A brow at 320×200 is about fourteen pixels wide and three tall. Its **shape**
|
||
carries almost nothing at that size; its **height above the eye** carries the
|
||
expression, and a brow raise is the most legible beat on a face. So the ring is
|
||
traced and the height is quantised — the same split the eyes got, where the lid
|
||
is a traced feature and the iris a quantised primitive.
|
||
|
||
The decomposition matters. The traced ring already contains the real height, so
|
||
adding a quantised raise on top would move the brow twice. Instead the height is
|
||
measured *out* of the ring, quantised, and put back: the shape that renders is
|
||
his, at a height that snaps between a few levels and holds.
|
||
|
||
Height is measured at **both ends**, not as one number, because raise and tilt
|
||
are different expressions out of one mechanism — both ends up is surprise, inner
|
||
up alone is worry, inner down is anger. They share a dwell, so the brow hits its
|
||
pose in one frame instead of crawling into it with one end arriving first.
|
||
|
||
It is measured against the eye's **corner midpoint**, never its lid — the same
|
||
trap the gaze origin has, and worth avoiding twice: brows and lids move together
|
||
constantly, so a brow that jumped on every blink would read as a tic. The rest
|
||
pose comes from the take **median**, not the neutral frame, for the same reason
|
||
gaze does: that frame is chosen by minimum mouth aperture and says nothing
|
||
whatever about the brows.
|
||
|
||
Two correspondences are resolved from geometry rather than declared: which ring
|
||
is which brow, and which end of a ring is the outer one. The second matters more
|
||
— get it backwards and the tilt mirrors, so worry renders as its own opposite,
|
||
which reads as a directed performance choice and would never be questioned.
|
||
Which *edge* of the brow is the upper one is deliberately left unresolved: it
|
||
traverses the same ring the other way round, an even-odd fill has no winding,
|
||
and the two ends still land on fixed slots either way.
|
||
|
||
| Knob | What it does |
|
||
| --- | --- |
|
||
| brow vertices | Ring vertex budget, off a 10-slot ring. |
|
||
| brow weight | Thickens the ring outward. It needs it at three pixels tall. |
|
||
| brow raise gain | Exaggerates or damps the raise. |
|
||
| brow step | The pixel grid the height snaps to. 0 = off. |
|
||
| brow dwell | How long a new height must hold. Shared across both ends. |
|
||
|
||
## The plate is reference, not art
|
||
|
||
The plate layer has several representations because its job changes. Cycle with
|
||
<kbd>B</kbd>:
|
||
|
||
| Mode | For |
|
||
| --- | --- |
|
||
| photo dim / photo | **Drawing over.** The source frame, stabilised. |
|
||
| posterized | The footage quantised into the ramp — a look test asked of the source rather than of a drawing. |
|
||
| oval | A flat stand-in, to judge the mouth against something. |
|
||
| oval + photo | Checking the stand-in against the real head. |
|
||
| none | Mouth alone. |
|
||
|
||
Photo modes are **registered**: the frame is mapped into raster space through the
|
||
same transform chain the contours go through, so the head sits still and a
|
||
drawing traced from it is already aligned to the mouth. An unregistered underlay
|
||
would be decorative.
|
||
|
||
`MediaPipe`'s face oval is the *face* boundary — it cuts at the hairline and
|
||
excludes hair, ears, jaw underside and neck — so as a head silhouette it is an
|
||
egg by construction, and no landmark precision fixes that. Hence the photo.
|
||
|
||
**Save frame 4x** writes the current registered composite as a 1280×800 PNG to
|
||
draw on.
|
||
|
||
## Paint — background cels
|
||
|
||
**A sketch.** It exists to test whether the aesthetic holds when a human draws
|
||
the background instead of the tracker deriving it, and it is meant to be
|
||
replaced by a real paint surface with onion skin and undo. It is one dependency-
|
||
free module, `js/paint.js`, so throwing it away is a delete rather than surgery.
|
||
|
||
Cels are drawn on the frames that get their own drawing and **hold until the
|
||
next one** — the same rule the plate follows, and literally the same lookup. You
|
||
can scrub anywhere and keep drawing on the cel you can see; the header says
|
||
which one you are editing and how far it holds.
|
||
|
||
- **pen** — click to place vertices, click the green box on the first one (or
|
||
<kbd>Enter</kbd> / double-click) to close.
|
||
- **edit** — click a shape to select, drag a vertex or the whole shape,
|
||
<kbd>Shift</kbd>-click an edge to insert a vertex, <kbd>Alt</kbd>-click one to
|
||
remove it, <kbd>Del</kbd> to delete the layer.
|
||
- **Layers** stack Photoshop-style, front at the top, with per-layer colour,
|
||
show/hide and reorder.
|
||
- **Drag a frame** from the strip onto the canvas to seed this cel from that
|
||
one, every layer, as a deep copy. **Copy previous** does the same for the
|
||
previous kept frame, which is the case you reach for constantly.
|
||
|
||
Two rules are enforced rather than left to discipline. Colours are **palette
|
||
indices**, so you cannot pick one that is not in the ramp — sampling colour from
|
||
the source is the one move `docs/design.md` says is irrecoverable. And vertices
|
||
**snap to the 320×200 grid**, because on a hard-edged indexed rasteriser a shape
|
||
nudged by 0.4px moves an edge by a whole pixel or not at all depending on where
|
||
it lands, which shimmers instead of holding.
|
||
|
||
Drawings autosave to `localStorage` per take name. They are the only thing in
|
||
the tool a person made by hand; everything else regenerates. They are not in the
|
||
`.take` export yet.
|
||
|
||
## Two kinds of sparseness
|
||
|
||
Sparseness has two unrelated causes, and conflating them was the original design
|
||
error here. **Aesthetic** sparseness is set by the extraction rate — pick 12fps and
|
||
you have already chosen your timing. **Labour** sparseness is a human drawing
|
||
each one, and it binds only on the plate.
|
||
|
||
Aesthetic sparseness is the **exposure** control, not the extraction rate —
|
||
making it a render-time grid means auditioning 12 against 24 costs a dropdown
|
||
instead of a re-rip and a full re-detection.
|
||
|
||
So the mouth keeps **every** frame: it is traced, and therefore free. In limited
|
||
animation lip sync is routinely the densest element, on 1s, while heads hold on
|
||
2s and 3s.
|
||
|
||
The frame strip is the editing surface for the other half: which frames need
|
||
their own plate drawing. Everything starts kept; delete what you don't want.
|
||
**Suggest** runs error-tolerance decimation over head pose as a starting point,
|
||
then you hand-correct.
|
||
|
||
**Mouth lead** is not a correction for a bug. A centred moving average has no
|
||
phase lag, so smoothing does not delay anything — but it blurs onsets, and the
|
||
visually salient moment of a mouth opening moves later even though the mean does
|
||
not. Animators also draw mouth shapes one or two frames ahead of the sound as
|
||
standard practice. Only the performance tracks shift; the head stays with the
|
||
audio, since it is the mouth that should anticipate. The lead is baked into the
|
||
exported take, so the renderer never needs to know about it.
|
||
|
||
`contour avg` is a deliberate, bounded exception to "never smooth the contour" in
|
||
`docs/design.md`. That rule held while keys were sparse, because sampling
|
||
at velocity minima rejected detector noise for free. With a key on every frame it
|
||
does not, so a radius shorter than the shortest articulation worth keeping is
|
||
justified — at 12fps, articulation spans 3–6 frames and detector noise is
|
||
per-frame, so ±1 separates them and ±3 starts eating speech.
|
||
|
||
## Tests
|
||
|
||
```sh
|
||
chromium --headless --virtual-time-budget=8000 --dump-dom \
|
||
http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+'
|
||
```
|
||
|
||
Or open `selftest.html`. 105 assertions over the stages below detection, plus a
|
||
wiring cross-check: every `el('id')` in `app.js` must exist in `index.html`. A
|
||
knob wired in one but not the other throws during wiring, which aborts the rest
|
||
of the module and leaves a blank page — a symptom that points nowhere near its
|
||
cause, and which has happened twice.
|
||
|
||
The ring-simplicity check is the load-bearing one. Because `hold` parts *cut*
|
||
between poses instead of interpolating, a ring whose vertex order is wrong
|
||
self-intersects and renders as blocks meeting at corners — and it is invisible at
|
||
odd `verts/2` and obvious at even, so it needs an assertion rather than an
|
||
eyeball.
|
||
|
||
## Not done yet
|
||
|
||
Hand-drawn head plates and per-plate mouth slots (the strip
|
||
decides *which frames need one*, but you cannot yet supply the drawing); real
|
||
performer→character calibration (currently identity, fitting the face oval to the
|
||
canvas); the override layer; anything on the Animator Pro side. The plate is a
|
||
face-oval polygon per kept frame — it exists so the mouth has a face to read
|
||
against, not to look good.
|