# arthur — design A small animation suite for turning live-action video into work that reads as hand-authored, in the idiom of *Another World* (Éric Chahi, 1991): flat shapes, a tiny palette, hard edges, motion carried by silhouette. ## The aesthetic is a representation, not a filter Chahi did not process video. He shot reference footage and hand-traced polygons over it in a custom editor; the engine stored and replayed polygon lists, never bitmaps. Three properties follow, and all three are load-bearing: 1. **Temporal identity.** The same shape, with the same vertex count and vertex order, *edited* across frames. A shape re-detected independently each frame produces a new contour every frame. That boils, and boiling reads as "filter" within half a second no matter how good the individual shapes are. 2. **Stylised timing.** Poses held, and for most parts a hard cut between them rather than an interpolation. A distinct pose on every frame of 24fps footage looks like video even when every pose is a polygon. 3. **Authored colour.** A small fixed ramp with two or three tones per part, chosen by a person. Sampling colour from the source produces a pixel-art filter immediately and irrecoverably. None of those are computer-vision problems. That is the whole argument about where CV belongs. ## The constraint is the point 320×200, indexed palette, flat fills, no antialiasing. That is inherited from Animator Pro, where this work started, but it is not an accident to be modernised away — it is why the output looks right. The rasteriser writes palette indices into a byte buffer and expands to RGBA only at the end, precisely so no canvas antialiasing can soften an edge. Modern conveniences belong in the *workflow* — instant feedback, audio, real undo, scrubbing, layers. Not in the output. ## CV tracks; it does not draw - **Its job.** Where the rigid features of the face are, what the head's transform is, where the lip contour runs, which frame is a motion extreme. - **Its non-job.** Deciding what the shapes are. A face mesh offers 468 points; an animator's mouth is eight. **That decimation ratio is the style.** It is a decision encoded in the tool, not a measurement extracted from footage. Every number crossing from analysis into rendering is a *parameter* or a *correspondence*. Every number determining how something looks comes from the artist or a fixed authored table. ## Two kinds of part | Kind | Source | Vocabulary | Interp | | --- | --- | --- | --- | | **Plate** — head, hair, body | Hand-drawn | Closed: a few drawings per character | hold | | **Feature** — mouth, lids | Rotoscoped from landmarks | Open: derived from this take | hold | | **Interior** — mouth interior, teeth | Image content within a feature | Open | hold | | **Primitive** — iris | Landmark centroid as a disc | Quantised | hold | The asymmetry is deliberate, and it is the opposite choice in each case. A **closed vocabulary is wrong for the mouth.** Pre-authored mouth shapes are never *this performance's* shapes, and that genericness is what the project exists to avoid. So the mouth gets an open vocabulary derived from footage, and the stylisation applies to its **timing** instead. A **closed vocabulary is right for the head.** The head carries structure, not performance. A handful of hand-drawn angles is better, because you drew them — and it means the head never needs segmentation, contour tracking, or boil avoidance. ## Two kinds of sparseness Conflating these was the original design error. **Aesthetic** sparseness is set by the extraction rate: pick 12fps and the timing is already chosen. **Labour** sparseness is a human drawing each one, and it binds only on the plate. So the mouth keeps **every** frame — it is traced, and therefore free. In limited animation lip sync is routinely the densest element, on 1s, while heads hold on 2s and 3s. The other half is frame removal: which frames need their own plate drawing. Everything starts kept; delete what you don't want. Automatic suggestion runs error-tolerance decimation over head pose as a starting point, then you hand-correct. ## Stabilisation A rotoscoped mouth is only reusable if expressed independently of where the head was. That inverts the obvious composition: ```text mouth_local(t) = anchor(t)⁻¹ · mouth_world(t) ``` - **Rigid landmarks only** for the fit: eye corners, nose bridge, nose tip. Including a feature that moves bleeds performance into the stabilisation. - **Similarity, not affine or homography.** The extra degrees of freedom absorb head rotation as distortion and smear it into the mouth. Four DOF removes exactly translation, roll and depth scale, leaving yaw and pitch as a measurable residual. - **Reference is the mean** configuration over the shot, not frame zero. - **Smooth the transform, never the contour** — with one bounded exception, below. - Landmarks convert to an **isotropic** space first (unit = one image height). MediaPipe normalises x by width and y by height, so its space is stretched; a "similarity" fitted there is not one. Out-of-plane rotation cannot be removed by any 2D transform. The answer is not a better transform: it is to draw the head at each angle and let each drawing declare where its mouth sits. Foreshortening becomes authored metadata. ## Hold versus interpolate Another World had no inbetweening engine: polygon sets played back frame by frame at a low rate. For parts that should read as hand-animated snaps — mouths above all — **cut between keys, do not interpolate.** A tweened mouth is rubbery and reads as puppet software immediately. **Fixed topology is non-negotiable for cut parts.** Vertex *meanings* must be stable across every key, so a closed mouth and a wide-open mouth are the same polygon at different positions. Interpolation partially hides ordering drift; cutting exposes it completely, as static. ## Deriving shapes from pixels without boiling Some things have no landmarks — teeth, tongue, anything inside the lips. The hazard is vertex correspondence: a traced contour reorders between frames. Two escapes, both used: - **Extract a scalar, not a shape**, where the shape can be derived from geometry you already trust. - **Radial sampling** where a real contour is needed. March outward from the blob's centroid along N fixed directions; vertex *k* is then always "the extent in direction *k*". Correspondence holds by construction, the count is fixed, temporal smoothing is well defined, and the star-shaped result suits flat colour. ## The bounded smoothing exception *Smooth the transform, never the contour* held while keys were sparse: sampling at velocity minima rejected per-frame detector noise for free. With a key on every frame it does not, so a short contour average is justified — the window must stay **shorter than the shortest articulation worth keeping**. At 12fps, mouth movement spans 3–6 frames and detector noise is per-frame, so ±1 separates them and ±3 starts eating speech. At 24fps, double it. ## Performance tracks may lead the clock A centred moving average has no phase lag, but it blurs onsets, so the visually salient moment of a mouth opening moves later even though the mean does not. Animators also draw mouth shapes a frame or two ahead of the sound as standard practice. Only performance tracks shift; the head stays with the audio, because it is the mouth that should anticipate. ## Colour discipline - A small ramp. Another World ran 16 colours at 320×200. - Two or three tones per part: base, shadow, occasionally a rim. - No dithering, no antialiasing anywhere. - Never sample colour from the source. Analysis may report *which* tone a region should be; it must never report an RGB value. - Plate art and generated parts share one palette, authored together. ## Three artifacts, not two | Artifact | Cost | Regenerated when | | --- | --- | --- | | Dense track — landmarks and fitted anchor per frame | Minutes, once per shot | Re-shoot, or a model change | | Take — parts, keys, kept frames | Instant | Every knob change | | Render | — | Continuous | Keep the dense track: re-keying is then a regenerate, never a re-trace. And because the take is disposable, **hand corrections must not live in it** — they belong in an override layer keyed by `(part, frame)`, applied on top at render time. ## Architecture The organising idea is an **indexed raster core**, a **part/track model**, and **pluggable sources** feeding it. Sources so far are rotoscope and image-derived interior; hand-drawn and procedural are the obvious next ones. Everything else — stabilisation, key selection, frame removal, palette — operates on the track model and is source-agnostic. Current modules: | Module | Role | | --- | --- | | `landmarks.js` | Index tables. Ring arrays are ordered traversals: slot position *is* vertex identity. | | `mathutil.js` | Similarity fit, Procrustes mean, temporal smoothing. | | `pipeline.js` | Stabilise → subsample → key-select → frame-removal. | | `interior.js` | Teeth from image content: Otsu, morphology, components, radial contour. | | `underlay.js` | Registered photo reference and palette posterisation. | | `raster.js` | Indexed scanline fill. No antialiasing, by construction. | | `take.js` | Take-file writer. | | `synth.js` | Synthetic landmarks, so everything below detection is testable without a video. | | `selftest.js` | Assertions, including source-level wiring checks. | ## Not built yet - **A paint surface.** The plates have nowhere to be drawn. This is the largest gap between "tool" and "suite": a pixel paint canvas with onion skin, palette constraint, and the registered underlay behind it. - Eyes, irises, brows as parts. - Plate libraries with per-plate mouth slots. - Real performer→character calibration (currently identity). - The override layer. - Export beyond the take file — FLI/FLC would connect back to the lineage and is a simple format. ## Lineage This began as an analysis front-end for Animator Pro, whose render path was the intended destination. `Animator-Pro-SDL/docs/roto-puppet.md` documents that design and what the Poco/FLX machinery can and cannot do. The reason for leaving is in that document's own rule — *never make a timing decision that requires a full render to evaluate* — which, once honoured, left the host with nothing to do but write the file.