220 lines
10 KiB
Markdown
220 lines
10 KiB
Markdown
|
|
# arthur — design
|
|||
|
|
|
|||
|
|
A small animation suite for turning live-action video into work that reads as
|
|||
|
|
hand-authored, in the idiom of *Another World* (Éric Chahi, 1991): flat shapes,
|
|||
|
|
a tiny palette, hard edges, motion carried by silhouette.
|
|||
|
|
|
|||
|
|
## The aesthetic is a representation, not a filter
|
|||
|
|
|
|||
|
|
Chahi did not process video. He shot reference footage and hand-traced polygons
|
|||
|
|
over it in a custom editor; the engine stored and replayed polygon lists, never
|
|||
|
|
bitmaps. Three properties follow, and all three are load-bearing:
|
|||
|
|
|
|||
|
|
1. **Temporal identity.** The same shape, with the same vertex count and vertex
|
|||
|
|
order, *edited* across frames. A shape re-detected independently each frame
|
|||
|
|
produces a new contour every frame. That boils, and boiling reads as "filter"
|
|||
|
|
within half a second no matter how good the individual shapes are.
|
|||
|
|
2. **Stylised timing.** Poses held, and for most parts a hard cut between them
|
|||
|
|
rather than an interpolation. A distinct pose on every frame of 24fps footage
|
|||
|
|
looks like video even when every pose is a polygon.
|
|||
|
|
3. **Authored colour.** A small fixed ramp with two or three tones per part,
|
|||
|
|
chosen by a person. Sampling colour from the source produces a pixel-art
|
|||
|
|
filter immediately and irrecoverably.
|
|||
|
|
|
|||
|
|
None of those are computer-vision problems. That is the whole argument about
|
|||
|
|
where CV belongs.
|
|||
|
|
|
|||
|
|
## The constraint is the point
|
|||
|
|
|
|||
|
|
320×200, indexed palette, flat fills, no antialiasing. That is inherited from
|
|||
|
|
Animator Pro, where this work started, but it is not an accident to be
|
|||
|
|
modernised away — it is why the output looks right. The rasteriser writes palette
|
|||
|
|
indices into a byte buffer and expands to RGBA only at the end, precisely so no
|
|||
|
|
canvas antialiasing can soften an edge.
|
|||
|
|
|
|||
|
|
Modern conveniences belong in the *workflow* — instant feedback, audio, real
|
|||
|
|
undo, scrubbing, layers. Not in the output.
|
|||
|
|
|
|||
|
|
## CV tracks; it does not draw
|
|||
|
|
|
|||
|
|
- **Its job.** Where the rigid features of the face are, what the head's
|
|||
|
|
transform is, where the lip contour runs, which frame is a motion extreme.
|
|||
|
|
- **Its non-job.** Deciding what the shapes are. A face mesh offers 468 points;
|
|||
|
|
an animator's mouth is eight. **That decimation ratio is the style.** It is a
|
|||
|
|
decision encoded in the tool, not a measurement extracted from footage.
|
|||
|
|
|
|||
|
|
Every number crossing from analysis into rendering is a *parameter* or a
|
|||
|
|
*correspondence*. Every number determining how something looks comes from the
|
|||
|
|
artist or a fixed authored table.
|
|||
|
|
|
|||
|
|
## Two kinds of part
|
|||
|
|
|
|||
|
|
| Kind | Source | Vocabulary | Interp |
|
|||
|
|
| --- | --- | --- | --- |
|
|||
|
|
| **Plate** — head, hair, body | Hand-drawn | Closed: a few drawings per character | hold |
|
|||
|
|
| **Feature** — mouth, lids | Rotoscoped from landmarks | Open: derived from this take | hold |
|
|||
|
|
| **Interior** — mouth interior, teeth | Image content within a feature | Open | hold |
|
|||
|
|
| **Primitive** — iris | Landmark centroid as a disc | Quantised | hold |
|
|||
|
|
|
|||
|
|
The asymmetry is deliberate, and it is the opposite choice in each case.
|
|||
|
|
|
|||
|
|
A **closed vocabulary is wrong for the mouth.** Pre-authored mouth shapes are
|
|||
|
|
never *this performance's* shapes, and that genericness is what the project
|
|||
|
|
exists to avoid. So the mouth gets an open vocabulary derived from footage, and
|
|||
|
|
the stylisation applies to its **timing** instead.
|
|||
|
|
|
|||
|
|
A **closed vocabulary is right for the head.** The head carries structure, not
|
|||
|
|
performance. A handful of hand-drawn angles is better, because you drew them —
|
|||
|
|
and it means the head never needs segmentation, contour tracking, or boil
|
|||
|
|
avoidance.
|
|||
|
|
|
|||
|
|
## Two kinds of sparseness
|
|||
|
|
|
|||
|
|
Conflating these was the original design error. **Aesthetic** sparseness is set
|
|||
|
|
by the extraction rate: pick 12fps and the timing is already chosen. **Labour**
|
|||
|
|
sparseness is a human drawing each one, and it binds only on the plate.
|
|||
|
|
|
|||
|
|
So the mouth keeps **every** frame — it is traced, and therefore free. In limited
|
|||
|
|
animation lip sync is routinely the densest element, on 1s, while heads hold on
|
|||
|
|
2s and 3s.
|
|||
|
|
|
|||
|
|
The other half is frame removal: which frames need their own plate drawing.
|
|||
|
|
Everything starts kept; delete what you don't want. Automatic suggestion runs
|
|||
|
|
error-tolerance decimation over head pose as a starting point, then you
|
|||
|
|
hand-correct.
|
|||
|
|
|
|||
|
|
## Stabilisation
|
|||
|
|
|
|||
|
|
A rotoscoped mouth is only reusable if expressed independently of where the head
|
|||
|
|
was. That inverts the obvious composition:
|
|||
|
|
|
|||
|
|
```text
|
|||
|
|
mouth_local(t) = anchor(t)⁻¹ · mouth_world(t)
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
- **Rigid landmarks only** for the fit: eye corners, nose bridge, nose tip.
|
|||
|
|
Including a feature that moves bleeds performance into the stabilisation.
|
|||
|
|
- **Similarity, not affine or homography.** The extra degrees of freedom absorb
|
|||
|
|
head rotation as distortion and smear it into the mouth. Four DOF removes
|
|||
|
|
exactly translation, roll and depth scale, leaving yaw and pitch as a
|
|||
|
|
measurable residual.
|
|||
|
|
- **Reference is the mean** configuration over the shot, not frame zero.
|
|||
|
|
- **Smooth the transform, never the contour** — with one bounded exception,
|
|||
|
|
below.
|
|||
|
|
- Landmarks convert to an **isotropic** space first (unit = one image height).
|
|||
|
|
MediaPipe normalises x by width and y by height, so its space is stretched; a
|
|||
|
|
"similarity" fitted there is not one.
|
|||
|
|
|
|||
|
|
Out-of-plane rotation cannot be removed by any 2D transform. The answer is not a
|
|||
|
|
better transform: it is to draw the head at each angle and let each drawing
|
|||
|
|
declare where its mouth sits. Foreshortening becomes authored metadata.
|
|||
|
|
|
|||
|
|
## Hold versus interpolate
|
|||
|
|
|
|||
|
|
Another World had no inbetweening engine: polygon sets played back frame by
|
|||
|
|
frame at a low rate. For parts that should read as hand-animated snaps — mouths
|
|||
|
|
above all — **cut between keys, do not interpolate.** A tweened mouth is rubbery
|
|||
|
|
and reads as puppet software immediately.
|
|||
|
|
|
|||
|
|
**Fixed topology is non-negotiable for cut parts.** Vertex *meanings* must be
|
|||
|
|
stable across every key, so a closed mouth and a wide-open mouth are the same
|
|||
|
|
polygon at different positions. Interpolation partially hides ordering drift;
|
|||
|
|
cutting exposes it completely, as static.
|
|||
|
|
|
|||
|
|
## Deriving shapes from pixels without boiling
|
|||
|
|
|
|||
|
|
Some things have no landmarks — teeth, tongue, anything inside the lips. The
|
|||
|
|
hazard is vertex correspondence: a traced contour reorders between frames.
|
|||
|
|
|
|||
|
|
Two escapes, both used:
|
|||
|
|
|
|||
|
|
- **Extract a scalar, not a shape**, where the shape can be derived from
|
|||
|
|
geometry you already trust.
|
|||
|
|
- **Radial sampling** where a real contour is needed. March outward from the
|
|||
|
|
blob's centroid along N fixed directions; vertex *k* is then always "the extent
|
|||
|
|
in direction *k*". Correspondence holds by construction, the count is fixed,
|
|||
|
|
temporal smoothing is well defined, and the star-shaped result suits flat
|
|||
|
|
colour.
|
|||
|
|
|
|||
|
|
## The bounded smoothing exception
|
|||
|
|
|
|||
|
|
*Smooth the transform, never the contour* held while keys were sparse: sampling
|
|||
|
|
at velocity minima rejected per-frame detector noise for free. With a key on
|
|||
|
|
every frame it does not, so a short contour average is justified — the window
|
|||
|
|
must stay **shorter than the shortest articulation worth keeping**. At 12fps,
|
|||
|
|
mouth movement spans 3–6 frames and detector noise is per-frame, so ±1 separates
|
|||
|
|
them and ±3 starts eating speech. At 24fps, double it.
|
|||
|
|
|
|||
|
|
## Performance tracks may lead the clock
|
|||
|
|
|
|||
|
|
A centred moving average has no phase lag, but it blurs onsets, so the visually
|
|||
|
|
salient moment of a mouth opening moves later even though the mean does not.
|
|||
|
|
Animators also draw mouth shapes a frame or two ahead of the sound as standard
|
|||
|
|
practice. Only performance tracks shift; the head stays with the audio, because
|
|||
|
|
it is the mouth that should anticipate.
|
|||
|
|
|
|||
|
|
## Colour discipline
|
|||
|
|
|
|||
|
|
- A small ramp. Another World ran 16 colours at 320×200.
|
|||
|
|
- Two or three tones per part: base, shadow, occasionally a rim.
|
|||
|
|
- No dithering, no antialiasing anywhere.
|
|||
|
|
- Never sample colour from the source. Analysis may report *which* tone a region
|
|||
|
|
should be; it must never report an RGB value.
|
|||
|
|
- Plate art and generated parts share one palette, authored together.
|
|||
|
|
|
|||
|
|
## Three artifacts, not two
|
|||
|
|
|
|||
|
|
| Artifact | Cost | Regenerated when |
|
|||
|
|
| --- | --- | --- |
|
|||
|
|
| Dense track — landmarks and fitted anchor per frame | Minutes, once per shot | Re-shoot, or a model change |
|
|||
|
|
| Take — parts, keys, kept frames | Instant | Every knob change |
|
|||
|
|
| Render | — | Continuous |
|
|||
|
|
|
|||
|
|
Keep the dense track: re-keying is then a regenerate, never a re-trace. And
|
|||
|
|
because the take is disposable, **hand corrections must not live in it** — they
|
|||
|
|
belong in an override layer keyed by `(part, frame)`, applied on top at render
|
|||
|
|
time.
|
|||
|
|
|
|||
|
|
## Architecture
|
|||
|
|
|
|||
|
|
The organising idea is an **indexed raster core**, a **part/track model**, and
|
|||
|
|
**pluggable sources** feeding it. Sources so far are rotoscope and image-derived
|
|||
|
|
interior; hand-drawn and procedural are the obvious next ones. Everything else —
|
|||
|
|
stabilisation, key selection, frame removal, palette — operates on the track
|
|||
|
|
model and is source-agnostic.
|
|||
|
|
|
|||
|
|
Current modules:
|
|||
|
|
|
|||
|
|
| Module | Role |
|
|||
|
|
| --- | --- |
|
|||
|
|
| `landmarks.js` | Index tables. Ring arrays are ordered traversals: slot position *is* vertex identity. |
|
|||
|
|
| `mathutil.js` | Similarity fit, Procrustes mean, temporal smoothing. |
|
|||
|
|
| `pipeline.js` | Stabilise → subsample → key-select → frame-removal. |
|
|||
|
|
| `interior.js` | Teeth from image content: Otsu, morphology, components, radial contour. |
|
|||
|
|
| `underlay.js` | Registered photo reference and palette posterisation. |
|
|||
|
|
| `raster.js` | Indexed scanline fill. No antialiasing, by construction. |
|
|||
|
|
| `take.js` | Take-file writer. |
|
|||
|
|
| `synth.js` | Synthetic landmarks, so everything below detection is testable without a video. |
|
|||
|
|
| `selftest.js` | Assertions, including source-level wiring checks. |
|
|||
|
|
|
|||
|
|
## Not built yet
|
|||
|
|
|
|||
|
|
- **A paint surface.** The plates have nowhere to be drawn. This is the largest
|
|||
|
|
gap between "tool" and "suite": a pixel paint canvas with onion skin, palette
|
|||
|
|
constraint, and the registered underlay behind it.
|
|||
|
|
- Eyes, irises, brows as parts.
|
|||
|
|
- Plate libraries with per-plate mouth slots.
|
|||
|
|
- Real performer→character calibration (currently identity).
|
|||
|
|
- The override layer.
|
|||
|
|
- Export beyond the take file — FLI/FLC would connect back to the lineage and is
|
|||
|
|
a simple format.
|
|||
|
|
|
|||
|
|
## Lineage
|
|||
|
|
|
|||
|
|
This began as an analysis front-end for Animator Pro, whose render path was the
|
|||
|
|
intended destination. `Animator-Pro-SDL/docs/roto-puppet.md` documents that
|
|||
|
|
design and what the Poco/FLX machinery can and cannot do. The reason for leaving
|
|||
|
|
is in that document's own rule — *never make a timing decision that requires a
|
|||
|
|
full render to evaluate* — which, once honoured, left the host with nothing to do
|
|||
|
|
but write the file.
|