arthur/docs/design.md

220 lines
10 KiB
Markdown
Raw Normal View History

# arthur — design
A small animation suite for turning live-action video into work that reads as
hand-authored, in the idiom of *Another World* (Éric Chahi, 1991): flat shapes,
a tiny palette, hard edges, motion carried by silhouette.
## The aesthetic is a representation, not a filter
Chahi did not process video. He shot reference footage and hand-traced polygons
over it in a custom editor; the engine stored and replayed polygon lists, never
bitmaps. Three properties follow, and all three are load-bearing:
1. **Temporal identity.** The same shape, with the same vertex count and vertex
order, *edited* across frames. A shape re-detected independently each frame
produces a new contour every frame. That boils, and boiling reads as "filter"
within half a second no matter how good the individual shapes are.
2. **Stylised timing.** Poses held, and for most parts a hard cut between them
rather than an interpolation. A distinct pose on every frame of 24fps footage
looks like video even when every pose is a polygon.
3. **Authored colour.** A small fixed ramp with two or three tones per part,
chosen by a person. Sampling colour from the source produces a pixel-art
filter immediately and irrecoverably.
None of those are computer-vision problems. That is the whole argument about
where CV belongs.
## The constraint is the point
320×200, indexed palette, flat fills, no antialiasing. That is inherited from
Animator Pro, where this work started, but it is not an accident to be
modernised away — it is why the output looks right. The rasteriser writes palette
indices into a byte buffer and expands to RGBA only at the end, precisely so no
canvas antialiasing can soften an edge.
Modern conveniences belong in the *workflow* — instant feedback, audio, real
undo, scrubbing, layers. Not in the output.
## CV tracks; it does not draw
- **Its job.** Where the rigid features of the face are, what the head's
transform is, where the lip contour runs, which frame is a motion extreme.
- **Its non-job.** Deciding what the shapes are. A face mesh offers 468 points;
an animator's mouth is eight. **That decimation ratio is the style.** It is a
decision encoded in the tool, not a measurement extracted from footage.
Every number crossing from analysis into rendering is a *parameter* or a
*correspondence*. Every number determining how something looks comes from the
artist or a fixed authored table.
## Two kinds of part
| Kind | Source | Vocabulary | Interp |
| --- | --- | --- | --- |
| **Plate** — head, hair, body | Hand-drawn | Closed: a few drawings per character | hold |
| **Feature** — mouth, lids | Rotoscoped from landmarks | Open: derived from this take | hold |
| **Interior** — mouth interior, teeth | Image content within a feature | Open | hold |
| **Primitive** — iris | Landmark centroid as a disc | Quantised | hold |
The asymmetry is deliberate, and it is the opposite choice in each case.
A **closed vocabulary is wrong for the mouth.** Pre-authored mouth shapes are
never *this performance's* shapes, and that genericness is what the project
exists to avoid. So the mouth gets an open vocabulary derived from footage, and
the stylisation applies to its **timing** instead.
A **closed vocabulary is right for the head.** The head carries structure, not
performance. A handful of hand-drawn angles is better, because you drew them —
and it means the head never needs segmentation, contour tracking, or boil
avoidance.
## Two kinds of sparseness
Conflating these was the original design error. **Aesthetic** sparseness is set
by the extraction rate: pick 12fps and the timing is already chosen. **Labour**
sparseness is a human drawing each one, and it binds only on the plate.
So the mouth keeps **every** frame — it is traced, and therefore free. In limited
animation lip sync is routinely the densest element, on 1s, while heads hold on
2s and 3s.
The other half is frame removal: which frames need their own plate drawing.
Everything starts kept; delete what you don't want. Automatic suggestion runs
error-tolerance decimation over head pose as a starting point, then you
hand-correct.
## Stabilisation
A rotoscoped mouth is only reusable if expressed independently of where the head
was. That inverts the obvious composition:
```text
mouth_local(t) = anchor(t)⁻¹ · mouth_world(t)
```
- **Rigid landmarks only** for the fit: eye corners, nose bridge, nose tip.
Including a feature that moves bleeds performance into the stabilisation.
- **Similarity, not affine or homography.** The extra degrees of freedom absorb
head rotation as distortion and smear it into the mouth. Four DOF removes
exactly translation, roll and depth scale, leaving yaw and pitch as a
measurable residual.
- **Reference is the mean** configuration over the shot, not frame zero.
- **Smooth the transform, never the contour** — with one bounded exception,
below.
- Landmarks convert to an **isotropic** space first (unit = one image height).
MediaPipe normalises x by width and y by height, so its space is stretched; a
"similarity" fitted there is not one.
Out-of-plane rotation cannot be removed by any 2D transform. The answer is not a
better transform: it is to draw the head at each angle and let each drawing
declare where its mouth sits. Foreshortening becomes authored metadata.
## Hold versus interpolate
Another World had no inbetweening engine: polygon sets played back frame by
frame at a low rate. For parts that should read as hand-animated snaps — mouths
above all — **cut between keys, do not interpolate.** A tweened mouth is rubbery
and reads as puppet software immediately.
**Fixed topology is non-negotiable for cut parts.** Vertex *meanings* must be
stable across every key, so a closed mouth and a wide-open mouth are the same
polygon at different positions. Interpolation partially hides ordering drift;
cutting exposes it completely, as static.
## Deriving shapes from pixels without boiling
Some things have no landmarks — teeth, tongue, anything inside the lips. The
hazard is vertex correspondence: a traced contour reorders between frames.
Two escapes, both used:
- **Extract a scalar, not a shape**, where the shape can be derived from
geometry you already trust.
- **Radial sampling** where a real contour is needed. March outward from the
blob's centroid along N fixed directions; vertex *k* is then always "the extent
in direction *k*". Correspondence holds by construction, the count is fixed,
temporal smoothing is well defined, and the star-shaped result suits flat
colour.
## The bounded smoothing exception
*Smooth the transform, never the contour* held while keys were sparse: sampling
at velocity minima rejected per-frame detector noise for free. With a key on
every frame it does not, so a short contour average is justified — the window
must stay **shorter than the shortest articulation worth keeping**. At 12fps,
mouth movement spans 3–6 frames and detector noise is per-frame, so ±1 separates
them and ±3 starts eating speech. At 24fps, double it.
## Performance tracks may lead the clock
A centred moving average has no phase lag, but it blurs onsets, so the visually
salient moment of a mouth opening moves later even though the mean does not.
Animators also draw mouth shapes a frame or two ahead of the sound as standard
practice. Only performance tracks shift; the head stays with the audio, because
it is the mouth that should anticipate.
## Colour discipline
- A small ramp. Another World ran 16 colours at 320×200.
- Two or three tones per part: base, shadow, occasionally a rim.
- No dithering, no antialiasing anywhere.
- Never sample colour from the source. Analysis may report *which* tone a region
should be; it must never report an RGB value.
- Plate art and generated parts share one palette, authored together.
## Three artifacts, not two
| Artifact | Cost | Regenerated when |
| --- | --- | --- |
| Dense track — landmarks and fitted anchor per frame | Minutes, once per shot | Re-shoot, or a model change |
| Take — parts, keys, kept frames | Instant | Every knob change |
| Render | — | Continuous |
Keep the dense track: re-keying is then a regenerate, never a re-trace. And
because the take is disposable, **hand corrections must not live in it** — they
belong in an override layer keyed by `(part, frame)`, applied on top at render
time.
## Architecture
The organising idea is an **indexed raster core**, a **part/track model**, and
**pluggable sources** feeding it. Sources so far are rotoscope and image-derived
interior; hand-drawn and procedural are the obvious next ones. Everything else —
stabilisation, key selection, frame removal, palette — operates on the track
model and is source-agnostic.
Current modules:
| Module | Role |
| --- | --- |
| `landmarks.js` | Index tables. Ring arrays are ordered traversals: slot position *is* vertex identity. |
| `mathutil.js` | Similarity fit, Procrustes mean, temporal smoothing. |
| `pipeline.js` | Stabilise → subsample → key-select → frame-removal. |
| `interior.js` | Teeth from image content: Otsu, morphology, components, radial contour. |
| `underlay.js` | Registered photo reference and palette posterisation. |
| `raster.js` | Indexed scanline fill. No antialiasing, by construction. |
| `take.js` | Take-file writer. |
| `synth.js` | Synthetic landmarks, so everything below detection is testable without a video. |
| `selftest.js` | Assertions, including source-level wiring checks. |
## Not built yet
- **A paint surface.** The plates have nowhere to be drawn. This is the largest
gap between "tool" and "suite": a pixel paint canvas with onion skin, palette
constraint, and the registered underlay behind it.
- Eyes, irises, brows as parts.
- Plate libraries with per-plate mouth slots.
- Real performer→character calibration (currently identity).
- The override layer.
- Export beyond the take file — FLI/FLC would connect back to the lineage and is
a simple format.
## Lineage
This began as an analysis front-end for Animator Pro, whose render path was the
intended destination. `Animator-Pro-SDL/docs/roto-puppet.md` documents that
design and what the Poco/FLX machinery can and cannot do. The reason for leaving
is in that document's own rule — *never make a timing decision that requires a
full render to evaluate* — which, once honoured, left the host with nothing to do
but write the file.