arthur/docs/design.md
Your Name f827757f7d Become arthur: a standalone suite, not an Animator Pro front-end
The test renderer turned out to be the product. Everything that decides how the
work looks - stabilisation, reduction, timing, frame removal, palette - already
happens here, and the flat indexed output already reads the way it should.

The reason to leave is in the original design's own rule: never make a timing
decision that requires a full render to evaluate. Honouring that moved every
judgement out of Animator Pro, which left the host doing nothing but writing a
file, in exchange for modal UI, minutes-long renders, one-level undo, FLX delta
invariants, a single tween state and a cel singleton.

What does NOT change is the constraint. 320x200, indexed palette, flat fills,
no antialiasing - inherited, but load-bearing rather than accidental. The
rasteriser writes palette indices and expands to RGBA only at the end precisely
so nothing can soften an edge. Modern conveniences belong in the workflow.

Adds docs/design.md: the principles, carried over without the Poco/FLX/cel
machinery, plus architecture and an honest list of what is missing - the
largest gap being that plates still have nowhere to be drawn.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:47:41 -04:00

10 KiB
Raw Blame History

arthur — design

A small animation suite for turning live-action video into work that reads as hand-authored, in the idiom of Another World (Éric Chahi, 1991): flat shapes, a tiny palette, hard edges, motion carried by silhouette.

The aesthetic is a representation, not a filter

Chahi did not process video. He shot reference footage and hand-traced polygons over it in a custom editor; the engine stored and replayed polygon lists, never bitmaps. Three properties follow, and all three are load-bearing:

  1. Temporal identity. The same shape, with the same vertex count and vertex order, edited across frames. A shape re-detected independently each frame produces a new contour every frame. That boils, and boiling reads as "filter" within half a second no matter how good the individual shapes are.
  2. Stylised timing. Poses held, and for most parts a hard cut between them rather than an interpolation. A distinct pose on every frame of 24fps footage looks like video even when every pose is a polygon.
  3. Authored colour. A small fixed ramp with two or three tones per part, chosen by a person. Sampling colour from the source produces a pixel-art filter immediately and irrecoverably.

None of those are computer-vision problems. That is the whole argument about where CV belongs.

The constraint is the point

320×200, indexed palette, flat fills, no antialiasing. That is inherited from Animator Pro, where this work started, but it is not an accident to be modernised away — it is why the output looks right. The rasteriser writes palette indices into a byte buffer and expands to RGBA only at the end, precisely so no canvas antialiasing can soften an edge.

Modern conveniences belong in the workflow — instant feedback, audio, real undo, scrubbing, layers. Not in the output.

CV tracks; it does not draw

  • Its job. Where the rigid features of the face are, what the head's transform is, where the lip contour runs, which frame is a motion extreme.
  • Its non-job. Deciding what the shapes are. A face mesh offers 468 points; an animator's mouth is eight. That decimation ratio is the style. It is a decision encoded in the tool, not a measurement extracted from footage.

Every number crossing from analysis into rendering is a parameter or a correspondence. Every number determining how something looks comes from the artist or a fixed authored table.

Two kinds of part

Kind Source Vocabulary Interp
Plate — head, hair, body Hand-drawn Closed: a few drawings per character hold
Feature — mouth, lids Rotoscoped from landmarks Open: derived from this take hold
Interior — mouth interior, teeth Image content within a feature Open hold
Primitive — iris Landmark centroid as a disc Quantised hold

The asymmetry is deliberate, and it is the opposite choice in each case.

A closed vocabulary is wrong for the mouth. Pre-authored mouth shapes are never this performance's shapes, and that genericness is what the project exists to avoid. So the mouth gets an open vocabulary derived from footage, and the stylisation applies to its timing instead.

A closed vocabulary is right for the head. The head carries structure, not performance. A handful of hand-drawn angles is better, because you drew them — and it means the head never needs segmentation, contour tracking, or boil avoidance.

Two kinds of sparseness

Conflating these was the original design error. Aesthetic sparseness is set by the extraction rate: pick 12fps and the timing is already chosen. Labour sparseness is a human drawing each one, and it binds only on the plate.

So the mouth keeps every frame — it is traced, and therefore free. In limited animation lip sync is routinely the densest element, on 1s, while heads hold on 2s and 3s.

The other half is frame removal: which frames need their own plate drawing. Everything starts kept; delete what you don't want. Automatic suggestion runs error-tolerance decimation over head pose as a starting point, then you hand-correct.

Stabilisation

A rotoscoped mouth is only reusable if expressed independently of where the head was. That inverts the obvious composition:

mouth_local(t) = anchor(t)⁻¹ · mouth_world(t)
  • Rigid landmarks only for the fit: eye corners, nose bridge, nose tip. Including a feature that moves bleeds performance into the stabilisation.
  • Similarity, not affine or homography. The extra degrees of freedom absorb head rotation as distortion and smear it into the mouth. Four DOF removes exactly translation, roll and depth scale, leaving yaw and pitch as a measurable residual.
  • Reference is the mean configuration over the shot, not frame zero.
  • Smooth the transform, never the contour — with one bounded exception, below.
  • Landmarks convert to an isotropic space first (unit = one image height). MediaPipe normalises x by width and y by height, so its space is stretched; a "similarity" fitted there is not one.

Out-of-plane rotation cannot be removed by any 2D transform. The answer is not a better transform: it is to draw the head at each angle and let each drawing declare where its mouth sits. Foreshortening becomes authored metadata.

Hold versus interpolate

Another World had no inbetweening engine: polygon sets played back frame by frame at a low rate. For parts that should read as hand-animated snaps — mouths above all — cut between keys, do not interpolate. A tweened mouth is rubbery and reads as puppet software immediately.

Fixed topology is non-negotiable for cut parts. Vertex meanings must be stable across every key, so a closed mouth and a wide-open mouth are the same polygon at different positions. Interpolation partially hides ordering drift; cutting exposes it completely, as static.

Deriving shapes from pixels without boiling

Some things have no landmarks — teeth, tongue, anything inside the lips. The hazard is vertex correspondence: a traced contour reorders between frames.

Two escapes, both used:

  • Extract a scalar, not a shape, where the shape can be derived from geometry you already trust.
  • Radial sampling where a real contour is needed. March outward from the blob's centroid along N fixed directions; vertex k is then always "the extent in direction k". Correspondence holds by construction, the count is fixed, temporal smoothing is well defined, and the star-shaped result suits flat colour.

The bounded smoothing exception

Smooth the transform, never the contour held while keys were sparse: sampling at velocity minima rejected per-frame detector noise for free. With a key on every frame it does not, so a short contour average is justified — the window must stay shorter than the shortest articulation worth keeping. At 12fps, mouth movement spans 3–6 frames and detector noise is per-frame, so ±1 separates them and ±3 starts eating speech. At 24fps, double it.

Performance tracks may lead the clock

A centred moving average has no phase lag, but it blurs onsets, so the visually salient moment of a mouth opening moves later even though the mean does not. Animators also draw mouth shapes a frame or two ahead of the sound as standard practice. Only performance tracks shift; the head stays with the audio, because it is the mouth that should anticipate.

Colour discipline

  • A small ramp. Another World ran 16 colours at 320×200.
  • Two or three tones per part: base, shadow, occasionally a rim.
  • No dithering, no antialiasing anywhere.
  • Never sample colour from the source. Analysis may report which tone a region should be; it must never report an RGB value.
  • Plate art and generated parts share one palette, authored together.

Three artifacts, not two

Artifact Cost Regenerated when
Dense track — landmarks and fitted anchor per frame Minutes, once per shot Re-shoot, or a model change
Take — parts, keys, kept frames Instant Every knob change
Render — Continuous

Keep the dense track: re-keying is then a regenerate, never a re-trace. And because the take is disposable, hand corrections must not live in it — they belong in an override layer keyed by (part, frame), applied on top at render time.

Architecture

The organising idea is an indexed raster core, a part/track model, and pluggable sources feeding it. Sources so far are rotoscope and image-derived interior; hand-drawn and procedural are the obvious next ones. Everything else — stabilisation, key selection, frame removal, palette — operates on the track model and is source-agnostic.

Current modules:

Module Role
landmarks.js Index tables. Ring arrays are ordered traversals: slot position is vertex identity.
mathutil.js Similarity fit, Procrustes mean, temporal smoothing.
pipeline.js Stabilise → subsample → key-select → frame-removal.
interior.js Teeth from image content: Otsu, morphology, components, radial contour.
underlay.js Registered photo reference and palette posterisation.
raster.js Indexed scanline fill. No antialiasing, by construction.
take.js Take-file writer.
synth.js Synthetic landmarks, so everything below detection is testable without a video.
selftest.js Assertions, including source-level wiring checks.

Not built yet

  • A paint surface. The plates have nowhere to be drawn. This is the largest gap between "tool" and "suite": a pixel paint canvas with onion skin, palette constraint, and the registered underlay behind it.
  • Eyes, irises, brows as parts.
  • Plate libraries with per-plate mouth slots.
  • Real performer→character calibration (currently identity).
  • The override layer.
  • Export beyond the take file — FLI/FLC would connect back to the lineage and is a simple format.

Lineage

This began as an analysis front-end for Animator Pro, whose render path was the intended destination. Animator-Pro-SDL/docs/roto-puppet.md documents that design and what the Poco/FLX machinery can and cannot do. The reason for leaving is in that document's own rule — never make a timing decision that requires a full render to evaluate — which, once honoured, left the host with nothing to do but write the file.