diff --git a/README.md b/README.md index e2a9174..e19e0ef 100644 --- a/README.md +++ b/README.md @@ -1,7 +1,7 @@ # roto Video → **take file** builder for the Animator Pro rotoscope/puppet pipeline -described in `../docs/roto-puppet.md`. +described in `docs/design.md`. This is the analysis and tuning half. It stabilises a face out of a clip, reduces the lip contour to a handful of vertices, selects sparse keys on motion extremes, @@ -144,7 +144,7 @@ audio, since it is the mouth that should anticipate. The lead is baked into the exported take, so the renderer never needs to know about it. `contour avg` is a deliberate, bounded exception to "never smooth the contour" in -`../docs/roto-puppet.md`. That rule held while keys were sparse, because sampling +`docs/design.md`. That rule held while keys were sparse, because sampling at velocity minima rejected detector noise for free. With a key on every frame it does not, so a radius shorter than the shortest articulation worth keeping is justified — at 12fps, articulation spans 3–6 frames and detector noise is diff --git a/docs/design.md b/docs/design.md new file mode 100644 index 0000000..a614389 --- /dev/null +++ b/docs/design.md @@ -0,0 +1,219 @@ +# arthur — design + +A small animation suite for turning live-action video into work that reads as +hand-authored, in the idiom of *Another World* (Éric Chahi, 1991): flat shapes, +a tiny palette, hard edges, motion carried by silhouette. + +## The aesthetic is a representation, not a filter + +Chahi did not process video. He shot reference footage and hand-traced polygons +over it in a custom editor; the engine stored and replayed polygon lists, never +bitmaps. Three properties follow, and all three are load-bearing: + +1. **Temporal identity.** The same shape, with the same vertex count and vertex + order, *edited* across frames. A shape re-detected independently each frame + produces a new contour every frame. That boils, and boiling reads as "filter" + within half a second no matter how good the individual shapes are. +2. **Stylised timing.** Poses held, and for most parts a hard cut between them + rather than an interpolation. A distinct pose on every frame of 24fps footage + looks like video even when every pose is a polygon. +3. **Authored colour.** A small fixed ramp with two or three tones per part, + chosen by a person. Sampling colour from the source produces a pixel-art + filter immediately and irrecoverably. + +None of those are computer-vision problems. That is the whole argument about +where CV belongs. + +## The constraint is the point + +320×200, indexed palette, flat fills, no antialiasing. That is inherited from +Animator Pro, where this work started, but it is not an accident to be +modernised away — it is why the output looks right. The rasteriser writes palette +indices into a byte buffer and expands to RGBA only at the end, precisely so no +canvas antialiasing can soften an edge. + +Modern conveniences belong in the *workflow* — instant feedback, audio, real +undo, scrubbing, layers. Not in the output. + +## CV tracks; it does not draw + +- **Its job.** Where the rigid features of the face are, what the head's + transform is, where the lip contour runs, which frame is a motion extreme. +- **Its non-job.** Deciding what the shapes are. A face mesh offers 468 points; + an animator's mouth is eight. **That decimation ratio is the style.** It is a + decision encoded in the tool, not a measurement extracted from footage. + +Every number crossing from analysis into rendering is a *parameter* or a +*correspondence*. Every number determining how something looks comes from the +artist or a fixed authored table. + +## Two kinds of part + +| Kind | Source | Vocabulary | Interp | +| --- | --- | --- | --- | +| **Plate** — head, hair, body | Hand-drawn | Closed: a few drawings per character | hold | +| **Feature** — mouth, lids | Rotoscoped from landmarks | Open: derived from this take | hold | +| **Interior** — mouth interior, teeth | Image content within a feature | Open | hold | +| **Primitive** — iris | Landmark centroid as a disc | Quantised | hold | + +The asymmetry is deliberate, and it is the opposite choice in each case. + +A **closed vocabulary is wrong for the mouth.** Pre-authored mouth shapes are +never *this performance's* shapes, and that genericness is what the project +exists to avoid. So the mouth gets an open vocabulary derived from footage, and +the stylisation applies to its **timing** instead. + +A **closed vocabulary is right for the head.** The head carries structure, not +performance. A handful of hand-drawn angles is better, because you drew them — +and it means the head never needs segmentation, contour tracking, or boil +avoidance. + +## Two kinds of sparseness + +Conflating these was the original design error. **Aesthetic** sparseness is set +by the extraction rate: pick 12fps and the timing is already chosen. **Labour** +sparseness is a human drawing each one, and it binds only on the plate. + +So the mouth keeps **every** frame — it is traced, and therefore free. In limited +animation lip sync is routinely the densest element, on 1s, while heads hold on +2s and 3s. + +The other half is frame removal: which frames need their own plate drawing. +Everything starts kept; delete what you don't want. Automatic suggestion runs +error-tolerance decimation over head pose as a starting point, then you +hand-correct. + +## Stabilisation + +A rotoscoped mouth is only reusable if expressed independently of where the head +was. That inverts the obvious composition: + +```text +mouth_local(t) = anchor(t)⁻¹ · mouth_world(t) +``` + +- **Rigid landmarks only** for the fit: eye corners, nose bridge, nose tip. + Including a feature that moves bleeds performance into the stabilisation. +- **Similarity, not affine or homography.** The extra degrees of freedom absorb + head rotation as distortion and smear it into the mouth. Four DOF removes + exactly translation, roll and depth scale, leaving yaw and pitch as a + measurable residual. +- **Reference is the mean** configuration over the shot, not frame zero. +- **Smooth the transform, never the contour** — with one bounded exception, + below. +- Landmarks convert to an **isotropic** space first (unit = one image height). + MediaPipe normalises x by width and y by height, so its space is stretched; a + "similarity" fitted there is not one. + +Out-of-plane rotation cannot be removed by any 2D transform. The answer is not a +better transform: it is to draw the head at each angle and let each drawing +declare where its mouth sits. Foreshortening becomes authored metadata. + +## Hold versus interpolate + +Another World had no inbetweening engine: polygon sets played back frame by +frame at a low rate. For parts that should read as hand-animated snaps — mouths +above all — **cut between keys, do not interpolate.** A tweened mouth is rubbery +and reads as puppet software immediately. + +**Fixed topology is non-negotiable for cut parts.** Vertex *meanings* must be +stable across every key, so a closed mouth and a wide-open mouth are the same +polygon at different positions. Interpolation partially hides ordering drift; +cutting exposes it completely, as static. + +## Deriving shapes from pixels without boiling + +Some things have no landmarks — teeth, tongue, anything inside the lips. The +hazard is vertex correspondence: a traced contour reorders between frames. + +Two escapes, both used: + +- **Extract a scalar, not a shape**, where the shape can be derived from + geometry you already trust. +- **Radial sampling** where a real contour is needed. March outward from the + blob's centroid along N fixed directions; vertex *k* is then always "the extent + in direction *k*". Correspondence holds by construction, the count is fixed, + temporal smoothing is well defined, and the star-shaped result suits flat + colour. + +## The bounded smoothing exception + +*Smooth the transform, never the contour* held while keys were sparse: sampling +at velocity minima rejected per-frame detector noise for free. With a key on +every frame it does not, so a short contour average is justified — the window +must stay **shorter than the shortest articulation worth keeping**. At 12fps, +mouth movement spans 3–6 frames and detector noise is per-frame, so ±1 separates +them and ±3 starts eating speech. At 24fps, double it. + +## Performance tracks may lead the clock + +A centred moving average has no phase lag, but it blurs onsets, so the visually +salient moment of a mouth opening moves later even though the mean does not. +Animators also draw mouth shapes a frame or two ahead of the sound as standard +practice. Only performance tracks shift; the head stays with the audio, because +it is the mouth that should anticipate. + +## Colour discipline + +- A small ramp. Another World ran 16 colours at 320×200. +- Two or three tones per part: base, shadow, occasionally a rim. +- No dithering, no antialiasing anywhere. +- Never sample colour from the source. Analysis may report *which* tone a region + should be; it must never report an RGB value. +- Plate art and generated parts share one palette, authored together. + +## Three artifacts, not two + +| Artifact | Cost | Regenerated when | +| --- | --- | --- | +| Dense track — landmarks and fitted anchor per frame | Minutes, once per shot | Re-shoot, or a model change | +| Take — parts, keys, kept frames | Instant | Every knob change | +| Render | — | Continuous | + +Keep the dense track: re-keying is then a regenerate, never a re-trace. And +because the take is disposable, **hand corrections must not live in it** — they +belong in an override layer keyed by `(part, frame)`, applied on top at render +time. + +## Architecture + +The organising idea is an **indexed raster core**, a **part/track model**, and +**pluggable sources** feeding it. Sources so far are rotoscope and image-derived +interior; hand-drawn and procedural are the obvious next ones. Everything else — +stabilisation, key selection, frame removal, palette — operates on the track +model and is source-agnostic. + +Current modules: + +| Module | Role | +| --- | --- | +| `landmarks.js` | Index tables. Ring arrays are ordered traversals: slot position *is* vertex identity. | +| `mathutil.js` | Similarity fit, Procrustes mean, temporal smoothing. | +| `pipeline.js` | Stabilise → subsample → key-select → frame-removal. | +| `interior.js` | Teeth from image content: Otsu, morphology, components, radial contour. | +| `underlay.js` | Registered photo reference and palette posterisation. | +| `raster.js` | Indexed scanline fill. No antialiasing, by construction. | +| `take.js` | Take-file writer. | +| `synth.js` | Synthetic landmarks, so everything below detection is testable without a video. | +| `selftest.js` | Assertions, including source-level wiring checks. | + +## Not built yet + +- **A paint surface.** The plates have nowhere to be drawn. This is the largest + gap between "tool" and "suite": a pixel paint canvas with onion skin, palette + constraint, and the registered underlay behind it. +- Eyes, irises, brows as parts. +- Plate libraries with per-plate mouth slots. +- Real performer→character calibration (currently identity). +- The override layer. +- Export beyond the take file — FLI/FLC would connect back to the lineage and is + a simple format. + +## Lineage + +This began as an analysis front-end for Animator Pro, whose render path was the +intended destination. `Animator-Pro-SDL/docs/roto-puppet.md` documents that +design and what the Poco/FLX machinery can and cannot do. The reason for leaving +is in that document's own rule — *never make a timing decision that requires a +full render to evaluate* — which, once honoured, left the host with nothing to do +but write the file. diff --git a/index.html b/index.html index b18a832..cb59dd6 100644 --- a/index.html +++ b/index.html @@ -3,7 +3,7 @@
-