diff --git a/README.md b/README.md index 85b29bc..9f1503f 100644 --- a/README.md +++ b/README.md @@ -1,16 +1,18 @@ -# roto +# arthur -Video → **take file** builder for the Animator Pro rotoscope/puppet pipeline -described in `../docs/roto-puppet.md`. +An animation suite for turning live-action video into work that reads as +hand-authored — flat shapes, a tiny palette, hard edges, motion carried by +silhouette, in the idiom of *Another World*. -This is the analysis and tuning half. It stabilises a face out of a clip, reduces -the lip contour to a handful of vertices, selects sparse keys on motion extremes, -and previews the result as flat indexed fills — so the *look and timing* can be -judged in seconds rather than through a minutes-long Animator Pro render. It -emits a `.take` file; nothing here touches Animator Pro. +It stabilises a face out of a clip, reduces the lip contour to a handful of +vertices, derives teeth from image content, lets you decide which frames need +their own hand-drawn head, and renders the result as flat indexed fills at +320×200 with no antialiasing. -All policy lives here. The take arrives at the renderer with its keys already -chosen. +The constraint is deliberate. 320×200 and an indexed palette are inherited from +Animator Pro, where this started, but they are why the output looks right — +modern conveniences belong in the workflow, not the output. See +[docs/design.md](docs/design.md). ## Run @@ -19,6 +21,9 @@ python3 -m http.server 8777 # from this directory # open http://127.0.0.1:8777 ``` +Static files and ES modules — no build step, no dependencies beyond MediaPipe's +wasm, which is fetched from a CDN on first use. + **Synthetic take** needs no video and exercises everything below detection. For real footage: @@ -57,11 +62,43 @@ is hand-drawn head plates, which this tool does not yet do. | Knob | What it does | | --- | --- | | vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. | +| mouth lead ±f | Shifts the performance tracks earlier against the audio and the head. `[` `]`. | | contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. | | anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform*. | | closed-mouth cut | Aperture below which the mouth interior is emitted as `hidden`. | | suggest tolerance | Max head movement before a new plate drawing is required. Affects **Suggest** only. | +## Teeth + +MediaPipe has no landmarks inside the lips: the inner ring bounds the cavity and +everything within it is just pixels. So teeth come from the image. + +The hazard is vertex correspondence — a traced contour reorders between frames +and boils. The way out for a blob specifically is **radial sampling**: march +outward from the blob's centroid along N fixed directions and take the last pixel +inside. Vertex *k* is then always "the extent in direction *k*", so correspondence +holds by construction, the vertex count is fixed, and temporal smoothing is well +defined with no reordering possible. It also produces a star-shaped reduction, +which is what flat blocks of colour want. + +The **teeth measurement** panel shows exactly what is sampled: region dimmed, +kept pixels green, extracted contour amber. Tune against that, not the numbers. + +| Knob | What it does | +| --- | --- | +| teeth contrast | Gate on the separation between the cavity's dark and bright class means. Otsu always returns *some* threshold, so this is what stops it inventing teeth in a dark mouth. | +| cavity erode | Pulls the sampled region in from the lip edge — MediaPipe's inner lip landmarks sit slightly outside the real opening, and lips are bright. | +| blob grow/erode | Resizes the found blob. An open pass always runs first to despeckle. | +| tongue reject | Drops pixels red relative to their own brightness. Teeth are near-neutral; tongue is not. | +| prefer upper | Biases component choice toward the top of the cavity. Area alone picks the tongue when the mouth is wide. | +| teeth vertices | Radial sample count. | +| teeth avg ±f | Temporal average over the contour. | +| teeth dwell | Frames a presence change must persist before it takes effect. | + +**Tongue** as its own part would work the same way, gated on redness instead of +brightness and biased low rather than high. Not implemented: it is not visible in +the test footage, which reads as a dark cavity with a bright upper-teeth band. + ## The plate is reference, not art The plate layer has several representations because its job changes. Cycle with @@ -103,8 +140,16 @@ their own plate drawing. Everything starts kept; delete what you don't want. **Suggest** runs error-tolerance decimation over head pose as a starting point, then you hand-correct. +**Mouth lead** is not a correction for a bug. A centred moving average has no +phase lag, so smoothing does not delay anything — but it blurs onsets, and the +visually salient moment of a mouth opening moves later even though the mean does +not. Animators also draw mouth shapes one or two frames ahead of the sound as +standard practice. Only the performance tracks shift; the head stays with the +audio, since it is the mouth that should anticipate. The lead is baked into the +exported take, so the renderer never needs to know about it. + `contour avg` is a deliberate, bounded exception to "never smooth the contour" in -`../docs/roto-puppet.md`. That rule held while keys were sparse, because sampling +`docs/design.md`. That rule held while keys were sparse, because sampling at velocity minima rejected detector noise for free. With a key on every frame it does not, so a radius shorter than the shortest articulation worth keeping is justified — at 12fps, articulation spans 3–6 frames and detector noise is @@ -117,7 +162,11 @@ chromium --headless --virtual-time-budget=8000 --dump-dom \ http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+' ``` -Or open `selftest.html`. 29 assertions over the stages below detection. +Or open `selftest.html`. 41 assertions over the stages below detection, plus a +wiring cross-check: every `el('id')` in `app.js` must exist in `index.html`. A +knob wired in one but not the other throws during wiring, which aborts the rest +of the module and leaves a blank page — a symptom that points nowhere near its +cause, and which has happened twice. The ring-simplicity check is the load-bearing one. Because `hold` parts *cut* between poses instead of interpolating, a ring whose vertex order is wrong diff --git a/docs/design.md b/docs/design.md new file mode 100644 index 0000000..a614389 --- /dev/null +++ b/docs/design.md @@ -0,0 +1,219 @@ +# arthur — design + +A small animation suite for turning live-action video into work that reads as +hand-authored, in the idiom of *Another World* (Éric Chahi, 1991): flat shapes, +a tiny palette, hard edges, motion carried by silhouette. + +## The aesthetic is a representation, not a filter + +Chahi did not process video. He shot reference footage and hand-traced polygons +over it in a custom editor; the engine stored and replayed polygon lists, never +bitmaps. Three properties follow, and all three are load-bearing: + +1. **Temporal identity.** The same shape, with the same vertex count and vertex + order, *edited* across frames. A shape re-detected independently each frame + produces a new contour every frame. That boils, and boiling reads as "filter" + within half a second no matter how good the individual shapes are. +2. **Stylised timing.** Poses held, and for most parts a hard cut between them + rather than an interpolation. A distinct pose on every frame of 24fps footage + looks like video even when every pose is a polygon. +3. **Authored colour.** A small fixed ramp with two or three tones per part, + chosen by a person. Sampling colour from the source produces a pixel-art + filter immediately and irrecoverably. + +None of those are computer-vision problems. That is the whole argument about +where CV belongs. + +## The constraint is the point + +320×200, indexed palette, flat fills, no antialiasing. That is inherited from +Animator Pro, where this work started, but it is not an accident to be +modernised away — it is why the output looks right. The rasteriser writes palette +indices into a byte buffer and expands to RGBA only at the end, precisely so no +canvas antialiasing can soften an edge. + +Modern conveniences belong in the *workflow* — instant feedback, audio, real +undo, scrubbing, layers. Not in the output. + +## CV tracks; it does not draw + +- **Its job.** Where the rigid features of the face are, what the head's + transform is, where the lip contour runs, which frame is a motion extreme. +- **Its non-job.** Deciding what the shapes are. A face mesh offers 468 points; + an animator's mouth is eight. **That decimation ratio is the style.** It is a + decision encoded in the tool, not a measurement extracted from footage. + +Every number crossing from analysis into rendering is a *parameter* or a +*correspondence*. Every number determining how something looks comes from the +artist or a fixed authored table. + +## Two kinds of part + +| Kind | Source | Vocabulary | Interp | +| --- | --- | --- | --- | +| **Plate** — head, hair, body | Hand-drawn | Closed: a few drawings per character | hold | +| **Feature** — mouth, lids | Rotoscoped from landmarks | Open: derived from this take | hold | +| **Interior** — mouth interior, teeth | Image content within a feature | Open | hold | +| **Primitive** — iris | Landmark centroid as a disc | Quantised | hold | + +The asymmetry is deliberate, and it is the opposite choice in each case. + +A **closed vocabulary is wrong for the mouth.** Pre-authored mouth shapes are +never *this performance's* shapes, and that genericness is what the project +exists to avoid. So the mouth gets an open vocabulary derived from footage, and +the stylisation applies to its **timing** instead. + +A **closed vocabulary is right for the head.** The head carries structure, not +performance. A handful of hand-drawn angles is better, because you drew them — +and it means the head never needs segmentation, contour tracking, or boil +avoidance. + +## Two kinds of sparseness + +Conflating these was the original design error. **Aesthetic** sparseness is set +by the extraction rate: pick 12fps and the timing is already chosen. **Labour** +sparseness is a human drawing each one, and it binds only on the plate. + +So the mouth keeps **every** frame — it is traced, and therefore free. In limited +animation lip sync is routinely the densest element, on 1s, while heads hold on +2s and 3s. + +The other half is frame removal: which frames need their own plate drawing. +Everything starts kept; delete what you don't want. Automatic suggestion runs +error-tolerance decimation over head pose as a starting point, then you +hand-correct. + +## Stabilisation + +A rotoscoped mouth is only reusable if expressed independently of where the head +was. That inverts the obvious composition: + +```text +mouth_local(t) = anchor(t)⁻¹ · mouth_world(t) +``` + +- **Rigid landmarks only** for the fit: eye corners, nose bridge, nose tip. + Including a feature that moves bleeds performance into the stabilisation. +- **Similarity, not affine or homography.** The extra degrees of freedom absorb + head rotation as distortion and smear it into the mouth. Four DOF removes + exactly translation, roll and depth scale, leaving yaw and pitch as a + measurable residual. +- **Reference is the mean** configuration over the shot, not frame zero. +- **Smooth the transform, never the contour** — with one bounded exception, + below. +- Landmarks convert to an **isotropic** space first (unit = one image height). + MediaPipe normalises x by width and y by height, so its space is stretched; a + "similarity" fitted there is not one. + +Out-of-plane rotation cannot be removed by any 2D transform. The answer is not a +better transform: it is to draw the head at each angle and let each drawing +declare where its mouth sits. Foreshortening becomes authored metadata. + +## Hold versus interpolate + +Another World had no inbetweening engine: polygon sets played back frame by +frame at a low rate. For parts that should read as hand-animated snaps — mouths +above all — **cut between keys, do not interpolate.** A tweened mouth is rubbery +and reads as puppet software immediately. + +**Fixed topology is non-negotiable for cut parts.** Vertex *meanings* must be +stable across every key, so a closed mouth and a wide-open mouth are the same +polygon at different positions. Interpolation partially hides ordering drift; +cutting exposes it completely, as static. + +## Deriving shapes from pixels without boiling + +Some things have no landmarks — teeth, tongue, anything inside the lips. The +hazard is vertex correspondence: a traced contour reorders between frames. + +Two escapes, both used: + +- **Extract a scalar, not a shape**, where the shape can be derived from + geometry you already trust. +- **Radial sampling** where a real contour is needed. March outward from the + blob's centroid along N fixed directions; vertex *k* is then always "the extent + in direction *k*". Correspondence holds by construction, the count is fixed, + temporal smoothing is well defined, and the star-shaped result suits flat + colour. + +## The bounded smoothing exception + +*Smooth the transform, never the contour* held while keys were sparse: sampling +at velocity minima rejected per-frame detector noise for free. With a key on +every frame it does not, so a short contour average is justified — the window +must stay **shorter than the shortest articulation worth keeping**. At 12fps, +mouth movement spans 3–6 frames and detector noise is per-frame, so ±1 separates +them and ±3 starts eating speech. At 24fps, double it. + +## Performance tracks may lead the clock + +A centred moving average has no phase lag, but it blurs onsets, so the visually +salient moment of a mouth opening moves later even though the mean does not. +Animators also draw mouth shapes a frame or two ahead of the sound as standard +practice. Only performance tracks shift; the head stays with the audio, because +it is the mouth that should anticipate. + +## Colour discipline + +- A small ramp. Another World ran 16 colours at 320×200. +- Two or three tones per part: base, shadow, occasionally a rim. +- No dithering, no antialiasing anywhere. +- Never sample colour from the source. Analysis may report *which* tone a region + should be; it must never report an RGB value. +- Plate art and generated parts share one palette, authored together. + +## Three artifacts, not two + +| Artifact | Cost | Regenerated when | +| --- | --- | --- | +| Dense track — landmarks and fitted anchor per frame | Minutes, once per shot | Re-shoot, or a model change | +| Take — parts, keys, kept frames | Instant | Every knob change | +| Render | — | Continuous | + +Keep the dense track: re-keying is then a regenerate, never a re-trace. And +because the take is disposable, **hand corrections must not live in it** — they +belong in an override layer keyed by `(part, frame)`, applied on top at render +time. + +## Architecture + +The organising idea is an **indexed raster core**, a **part/track model**, and +**pluggable sources** feeding it. Sources so far are rotoscope and image-derived +interior; hand-drawn and procedural are the obvious next ones. Everything else — +stabilisation, key selection, frame removal, palette — operates on the track +model and is source-agnostic. + +Current modules: + +| Module | Role | +| --- | --- | +| `landmarks.js` | Index tables. Ring arrays are ordered traversals: slot position *is* vertex identity. | +| `mathutil.js` | Similarity fit, Procrustes mean, temporal smoothing. | +| `pipeline.js` | Stabilise → subsample → key-select → frame-removal. | +| `interior.js` | Teeth from image content: Otsu, morphology, components, radial contour. | +| `underlay.js` | Registered photo reference and palette posterisation. | +| `raster.js` | Indexed scanline fill. No antialiasing, by construction. | +| `take.js` | Take-file writer. | +| `synth.js` | Synthetic landmarks, so everything below detection is testable without a video. | +| `selftest.js` | Assertions, including source-level wiring checks. | + +## Not built yet + +- **A paint surface.** The plates have nowhere to be drawn. This is the largest + gap between "tool" and "suite": a pixel paint canvas with onion skin, palette + constraint, and the registered underlay behind it. +- Eyes, irises, brows as parts. +- Plate libraries with per-plate mouth slots. +- Real performer→character calibration (currently identity). +- The override layer. +- Export beyond the take file — FLI/FLC would connect back to the lineage and is + a simple format. + +## Lineage + +This began as an analysis front-end for Animator Pro, whose render path was the +intended destination. `Animator-Pro-SDL/docs/roto-puppet.md` documents that +design and what the Poco/FLX machinery can and cannot do. The reason for leaving +is in that document's own rule — *never make a timing decision that requires a +full render to evaluate* — which, once honoured, left the host with nothing to do +but write the file. diff --git a/gl.html b/gl.html deleted file mode 100644 index b2b042c..0000000 --- a/gl.html +++ /dev/null @@ -1,6 +0,0 @@ -
?diff --git a/index.html b/index.html index 37475e3..cb59dd6 100644 --- a/index.html +++ b/index.html @@ -3,7 +3,7 @@ -