diff --git a/README.md b/README.md index 9f1503f..85b29bc 100644 --- a/README.md +++ b/README.md @@ -1,18 +1,16 @@ -# arthur +# roto -An animation suite for turning live-action video into work that reads as -hand-authored — flat shapes, a tiny palette, hard edges, motion carried by -silhouette, in the idiom of *Another World*. +Video → **take file** builder for the Animator Pro rotoscope/puppet pipeline +described in `../docs/roto-puppet.md`. -It stabilises a face out of a clip, reduces the lip contour to a handful of -vertices, derives teeth from image content, lets you decide which frames need -their own hand-drawn head, and renders the result as flat indexed fills at -320×200 with no antialiasing. +This is the analysis and tuning half. It stabilises a face out of a clip, reduces +the lip contour to a handful of vertices, selects sparse keys on motion extremes, +and previews the result as flat indexed fills — so the *look and timing* can be +judged in seconds rather than through a minutes-long Animator Pro render. It +emits a `.take` file; nothing here touches Animator Pro. -The constraint is deliberate. 320×200 and an indexed palette are inherited from -Animator Pro, where this started, but they are why the output looks right — -modern conveniences belong in the workflow, not the output. See -[docs/design.md](docs/design.md). +All policy lives here. The take arrives at the renderer with its keys already +chosen. ## Run @@ -21,9 +19,6 @@ python3 -m http.server 8777 # from this directory # open http://127.0.0.1:8777 ``` -Static files and ES modules — no build step, no dependencies beyond MediaPipe's -wasm, which is fetched from a CDN on first use. - **Synthetic take** needs no video and exercises everything below detection. For real footage: @@ -62,43 +57,11 @@ is hand-drawn head plates, which this tool does not yet do. | Knob | What it does | | --- | --- | | vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. | -| mouth lead ±f | Shifts the performance tracks earlier against the audio and the head. `[` `]`. | | contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. | | anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform*. | | closed-mouth cut | Aperture below which the mouth interior is emitted as `hidden`. | | suggest tolerance | Max head movement before a new plate drawing is required. Affects **Suggest** only. | -## Teeth - -MediaPipe has no landmarks inside the lips: the inner ring bounds the cavity and -everything within it is just pixels. So teeth come from the image. - -The hazard is vertex correspondence — a traced contour reorders between frames -and boils. The way out for a blob specifically is **radial sampling**: march -outward from the blob's centroid along N fixed directions and take the last pixel -inside. Vertex *k* is then always "the extent in direction *k*", so correspondence -holds by construction, the vertex count is fixed, and temporal smoothing is well -defined with no reordering possible. It also produces a star-shaped reduction, -which is what flat blocks of colour want. - -The **teeth measurement** panel shows exactly what is sampled: region dimmed, -kept pixels green, extracted contour amber. Tune against that, not the numbers. - -| Knob | What it does | -| --- | --- | -| teeth contrast | Gate on the separation between the cavity's dark and bright class means. Otsu always returns *some* threshold, so this is what stops it inventing teeth in a dark mouth. | -| cavity erode | Pulls the sampled region in from the lip edge — MediaPipe's inner lip landmarks sit slightly outside the real opening, and lips are bright. | -| blob grow/erode | Resizes the found blob. An open pass always runs first to despeckle. | -| tongue reject | Drops pixels red relative to their own brightness. Teeth are near-neutral; tongue is not. | -| prefer upper | Biases component choice toward the top of the cavity. Area alone picks the tongue when the mouth is wide. | -| teeth vertices | Radial sample count. | -| teeth avg ±f | Temporal average over the contour. | -| teeth dwell | Frames a presence change must persist before it takes effect. | - -**Tongue** as its own part would work the same way, gated on redness instead of -brightness and biased low rather than high. Not implemented: it is not visible in -the test footage, which reads as a dark cavity with a bright upper-teeth band. - ## The plate is reference, not art The plate layer has several representations because its job changes. Cycle with @@ -140,16 +103,8 @@ their own plate drawing. Everything starts kept; delete what you don't want. **Suggest** runs error-tolerance decimation over head pose as a starting point, then you hand-correct. -**Mouth lead** is not a correction for a bug. A centred moving average has no -phase lag, so smoothing does not delay anything — but it blurs onsets, and the -visually salient moment of a mouth opening moves later even though the mean does -not. Animators also draw mouth shapes one or two frames ahead of the sound as -standard practice. Only the performance tracks shift; the head stays with the -audio, since it is the mouth that should anticipate. The lead is baked into the -exported take, so the renderer never needs to know about it. - `contour avg` is a deliberate, bounded exception to "never smooth the contour" in -`docs/design.md`. That rule held while keys were sparse, because sampling +`../docs/roto-puppet.md`. That rule held while keys were sparse, because sampling at velocity minima rejected detector noise for free. With a key on every frame it does not, so a radius shorter than the shortest articulation worth keeping is justified — at 12fps, articulation spans 3–6 frames and detector noise is @@ -162,11 +117,7 @@ chromium --headless --virtual-time-budget=8000 --dump-dom \ http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+' ``` -Or open `selftest.html`. 41 assertions over the stages below detection, plus a -wiring cross-check: every `el('id')` in `app.js` must exist in `index.html`. A -knob wired in one but not the other throws during wiring, which aborts the rest -of the module and leaves a blank page — a symptom that points nowhere near its -cause, and which has happened twice. +Or open `selftest.html`. 29 assertions over the stages below detection. The ring-simplicity check is the load-bearing one. Because `hold` parts *cut* between poses instead of interpolating, a ring whose vertex order is wrong diff --git a/docs/design.md b/docs/design.md deleted file mode 100644 index a614389..0000000 --- a/docs/design.md +++ /dev/null @@ -1,219 +0,0 @@ -# arthur — design - -A small animation suite for turning live-action video into work that reads as -hand-authored, in the idiom of *Another World* (Éric Chahi, 1991): flat shapes, -a tiny palette, hard edges, motion carried by silhouette. - -## The aesthetic is a representation, not a filter - -Chahi did not process video. He shot reference footage and hand-traced polygons -over it in a custom editor; the engine stored and replayed polygon lists, never -bitmaps. Three properties follow, and all three are load-bearing: - -1. **Temporal identity.** The same shape, with the same vertex count and vertex - order, *edited* across frames. A shape re-detected independently each frame - produces a new contour every frame. That boils, and boiling reads as "filter" - within half a second no matter how good the individual shapes are. -2. **Stylised timing.** Poses held, and for most parts a hard cut between them - rather than an interpolation. A distinct pose on every frame of 24fps footage - looks like video even when every pose is a polygon. -3. **Authored colour.** A small fixed ramp with two or three tones per part, - chosen by a person. Sampling colour from the source produces a pixel-art - filter immediately and irrecoverably. - -None of those are computer-vision problems. That is the whole argument about -where CV belongs. - -## The constraint is the point - -320×200, indexed palette, flat fills, no antialiasing. That is inherited from -Animator Pro, where this work started, but it is not an accident to be -modernised away — it is why the output looks right. The rasteriser writes palette -indices into a byte buffer and expands to RGBA only at the end, precisely so no -canvas antialiasing can soften an edge. - -Modern conveniences belong in the *workflow* — instant feedback, audio, real -undo, scrubbing, layers. Not in the output. - -## CV tracks; it does not draw - -- **Its job.** Where the rigid features of the face are, what the head's - transform is, where the lip contour runs, which frame is a motion extreme. -- **Its non-job.** Deciding what the shapes are. A face mesh offers 468 points; - an animator's mouth is eight. **That decimation ratio is the style.** It is a - decision encoded in the tool, not a measurement extracted from footage. - -Every number crossing from analysis into rendering is a *parameter* or a -*correspondence*. Every number determining how something looks comes from the -artist or a fixed authored table. - -## Two kinds of part - -| Kind | Source | Vocabulary | Interp | -| --- | --- | --- | --- | -| **Plate** — head, hair, body | Hand-drawn | Closed: a few drawings per character | hold | -| **Feature** — mouth, lids | Rotoscoped from landmarks | Open: derived from this take | hold | -| **Interior** — mouth interior, teeth | Image content within a feature | Open | hold | -| **Primitive** — iris | Landmark centroid as a disc | Quantised | hold | - -The asymmetry is deliberate, and it is the opposite choice in each case. - -A **closed vocabulary is wrong for the mouth.** Pre-authored mouth shapes are -never *this performance's* shapes, and that genericness is what the project -exists to avoid. So the mouth gets an open vocabulary derived from footage, and -the stylisation applies to its **timing** instead. - -A **closed vocabulary is right for the head.** The head carries structure, not -performance. A handful of hand-drawn angles is better, because you drew them — -and it means the head never needs segmentation, contour tracking, or boil -avoidance. - -## Two kinds of sparseness - -Conflating these was the original design error. **Aesthetic** sparseness is set -by the extraction rate: pick 12fps and the timing is already chosen. **Labour** -sparseness is a human drawing each one, and it binds only on the plate. - -So the mouth keeps **every** frame — it is traced, and therefore free. In limited -animation lip sync is routinely the densest element, on 1s, while heads hold on -2s and 3s. - -The other half is frame removal: which frames need their own plate drawing. -Everything starts kept; delete what you don't want. Automatic suggestion runs -error-tolerance decimation over head pose as a starting point, then you -hand-correct. - -## Stabilisation - -A rotoscoped mouth is only reusable if expressed independently of where the head -was. That inverts the obvious composition: - -```text -mouth_local(t) = anchor(t)⁻¹ · mouth_world(t) -``` - -- **Rigid landmarks only** for the fit: eye corners, nose bridge, nose tip. - Including a feature that moves bleeds performance into the stabilisation. -- **Similarity, not affine or homography.** The extra degrees of freedom absorb - head rotation as distortion and smear it into the mouth. Four DOF removes - exactly translation, roll and depth scale, leaving yaw and pitch as a - measurable residual. -- **Reference is the mean** configuration over the shot, not frame zero. -- **Smooth the transform, never the contour** — with one bounded exception, - below. -- Landmarks convert to an **isotropic** space first (unit = one image height). - MediaPipe normalises x by width and y by height, so its space is stretched; a - "similarity" fitted there is not one. - -Out-of-plane rotation cannot be removed by any 2D transform. The answer is not a -better transform: it is to draw the head at each angle and let each drawing -declare where its mouth sits. Foreshortening becomes authored metadata. - -## Hold versus interpolate - -Another World had no inbetweening engine: polygon sets played back frame by -frame at a low rate. For parts that should read as hand-animated snaps — mouths -above all — **cut between keys, do not interpolate.** A tweened mouth is rubbery -and reads as puppet software immediately. - -**Fixed topology is non-negotiable for cut parts.** Vertex *meanings* must be -stable across every key, so a closed mouth and a wide-open mouth are the same -polygon at different positions. Interpolation partially hides ordering drift; -cutting exposes it completely, as static. - -## Deriving shapes from pixels without boiling - -Some things have no landmarks — teeth, tongue, anything inside the lips. The -hazard is vertex correspondence: a traced contour reorders between frames. - -Two escapes, both used: - -- **Extract a scalar, not a shape**, where the shape can be derived from - geometry you already trust. -- **Radial sampling** where a real contour is needed. March outward from the - blob's centroid along N fixed directions; vertex *k* is then always "the extent - in direction *k*". Correspondence holds by construction, the count is fixed, - temporal smoothing is well defined, and the star-shaped result suits flat - colour. - -## The bounded smoothing exception - -*Smooth the transform, never the contour* held while keys were sparse: sampling -at velocity minima rejected per-frame detector noise for free. With a key on -every frame it does not, so a short contour average is justified — the window -must stay **shorter than the shortest articulation worth keeping**. At 12fps, -mouth movement spans 3–6 frames and detector noise is per-frame, so ±1 separates -them and ±3 starts eating speech. At 24fps, double it. - -## Performance tracks may lead the clock - -A centred moving average has no phase lag, but it blurs onsets, so the visually -salient moment of a mouth opening moves later even though the mean does not. -Animators also draw mouth shapes a frame or two ahead of the sound as standard -practice. Only performance tracks shift; the head stays with the audio, because -it is the mouth that should anticipate. - -## Colour discipline - -- A small ramp. Another World ran 16 colours at 320×200. -- Two or three tones per part: base, shadow, occasionally a rim. -- No dithering, no antialiasing anywhere. -- Never sample colour from the source. Analysis may report *which* tone a region - should be; it must never report an RGB value. -- Plate art and generated parts share one palette, authored together. - -## Three artifacts, not two - -| Artifact | Cost | Regenerated when | -| --- | --- | --- | -| Dense track — landmarks and fitted anchor per frame | Minutes, once per shot | Re-shoot, or a model change | -| Take — parts, keys, kept frames | Instant | Every knob change | -| Render | — | Continuous | - -Keep the dense track: re-keying is then a regenerate, never a re-trace. And -because the take is disposable, **hand corrections must not live in it** — they -belong in an override layer keyed by `(part, frame)`, applied on top at render -time. - -## Architecture - -The organising idea is an **indexed raster core**, a **part/track model**, and -**pluggable sources** feeding it. Sources so far are rotoscope and image-derived -interior; hand-drawn and procedural are the obvious next ones. Everything else — -stabilisation, key selection, frame removal, palette — operates on the track -model and is source-agnostic. - -Current modules: - -| Module | Role | -| --- | --- | -| `landmarks.js` | Index tables. Ring arrays are ordered traversals: slot position *is* vertex identity. | -| `mathutil.js` | Similarity fit, Procrustes mean, temporal smoothing. | -| `pipeline.js` | Stabilise → subsample → key-select → frame-removal. | -| `interior.js` | Teeth from image content: Otsu, morphology, components, radial contour. | -| `underlay.js` | Registered photo reference and palette posterisation. | -| `raster.js` | Indexed scanline fill. No antialiasing, by construction. | -| `take.js` | Take-file writer. | -| `synth.js` | Synthetic landmarks, so everything below detection is testable without a video. | -| `selftest.js` | Assertions, including source-level wiring checks. | - -## Not built yet - -- **A paint surface.** The plates have nowhere to be drawn. This is the largest - gap between "tool" and "suite": a pixel paint canvas with onion skin, palette - constraint, and the registered underlay behind it. -- Eyes, irises, brows as parts. -- Plate libraries with per-plate mouth slots. -- Real performer→character calibration (currently identity). -- The override layer. -- Export beyond the take file — FLI/FLC would connect back to the lineage and is - a simple format. - -## Lineage - -This began as an analysis front-end for Animator Pro, whose render path was the -intended destination. `Animator-Pro-SDL/docs/roto-puppet.md` documents that -design and what the Poco/FLX machinery can and cannot do. The reason for leaving -is in that document's own rule — *never make a timing decision that requires a -full render to evaluate* — which, once honoured, left the host with nothing to do -but write the file. diff --git a/gl.html b/gl.html new file mode 100644 index 0000000..b2b042c --- /dev/null +++ b/gl.html @@ -0,0 +1,6 @@ +
?diff --git a/index.html b/index.html index cb59dd6..37475e3 100644 --- a/index.html +++ b/index.html @@ -3,7 +3,7 @@ -