arthur/README.md

453 lines
24 KiB
Markdown
Raw Permalink Normal View History

2026-09-24 15:49:05 -04:00
# arthur
2026-09-24 15:49:05 -04:00
An animation suite for turning live-action video into work that reads as
hand-authored — flat shapes, a tiny palette, hard edges, motion carried by
silhouette, in the idiom of *Another World*.
2026-09-24 15:49:05 -04:00
It stabilises a face out of a clip, reduces the lip contour to a handful of
vertices, derives teeth from image content, lets you decide which frames need
their own hand-drawn head, and renders the result as flat indexed fills at
320×200 with no antialiasing.
2026-09-24 15:49:05 -04:00
The constraint is deliberate. 320×200 and an indexed palette are inherited from
Animator Pro, where this started, but they are why the output looks right —
modern conveniences belong in the workflow, not the output. See
[docs/design.md](docs/design.md).
## ClojureScript port
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
The active port plays the synthetic take, accepts video uploads, transcodes them
to an H.264 proxy and decodable stream plus audio and tracing stills, analyzes real footage
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
for mouth, eyes, brows and pixel-derived teeth, and saves the project with
reusable analysis data. The step 8 data model
represents persistent feature IDs, eye pairs and feature-level observation gaps;
its controls are still pending. See the [port plan](docs/port-plan.md).
Serve the document from a Django backend, split into three tiers Step 9. The tier split was the work; Django was the easy half. Tier 1 — the authored scene — is the document, and it is addressed as independently versioned leaves rather than saved whole, so one vertex drag cannot clobber a collaborator's keying. `domain/leaf` is the document as path -> value; `domain/wire` puts it on the wire as transit, because JSON has neither integer map keys nor keywords and a save would quietly turn `{0 v}` into `{"0" v}`. Tier 2 — the dense channel blocks — is content-addressed by a hash over every input, with the detector version inside every key through the analysis the block descriptor names. `flow/address`'s `block-knobs` is the invalidation table, and `address-test` does not trust it: it re-freezes the take once per knob and asserts the biconditional, that a block's bytes changed if and only if its key changed. That found `brow-pos` not depending on `contour-avg` — the brow ring is smoothed, the raise is not. Tier 3 — frames and audio — is served by the hash of its bytes out of the same store. A manifest now names frames and carries a URL for each, so the frame layout stopped being a shared secret between a shell script and a ClojureScript namespace, and the `?v=` cache-buster went with it: a blob's name is the hash of its contents, so a stale copy is not a thing that can happen. The synthetic take's `audio.wav` moved to `static/arthur/` — an asset the project owns, not an extraction that churns. The server verifies rather than trusting a name it was handed: it recomputes every key from the descriptor stored beside it, refuses an analysis that declares no detector version, and refuses a document naming blocks it does not hold. It hashes the descriptor TEXT, because JS prints an integral double as `1` and Python as `1.0`, and a scheme where both ends re-render the numbers disagrees on the first parameter that happens to be whole. Two loose ends from step 8 closed on the way. `pack` no longer takes a `(track, frame)` predicate whose call sites each re-derived a feature from an index — every track names the feature it follows, which deleted five hand-maintained mappings. And `:dev-http` is gone: Django serves the page, shadow-cljs only builds into the staticfiles tree. 227 CLJS tests, 31 Django tests, green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 01:11:41 -04:00
```sh
mise install # both halves
pip install -r requirements.txt
mise exec -- python manage.py migrate
./do start # Django + frontend watcher
Serve the document from a Django backend, split into three tiers Step 9. The tier split was the work; Django was the easy half. Tier 1 — the authored scene — is the document, and it is addressed as independently versioned leaves rather than saved whole, so one vertex drag cannot clobber a collaborator's keying. `domain/leaf` is the document as path -> value; `domain/wire` puts it on the wire as transit, because JSON has neither integer map keys nor keywords and a save would quietly turn `{0 v}` into `{"0" v}`. Tier 2 — the dense channel blocks — is content-addressed by a hash over every input, with the detector version inside every key through the analysis the block descriptor names. `flow/address`'s `block-knobs` is the invalidation table, and `address-test` does not trust it: it re-freezes the take once per knob and asserts the biconditional, that a block's bytes changed if and only if its key changed. That found `brow-pos` not depending on `contour-avg` — the brow ring is smoothed, the raise is not. Tier 3 — frames and audio — is served by the hash of its bytes out of the same store. A manifest now names frames and carries a URL for each, so the frame layout stopped being a shared secret between a shell script and a ClojureScript namespace, and the `?v=` cache-buster went with it: a blob's name is the hash of its contents, so a stale copy is not a thing that can happen. The synthetic take's `audio.wav` moved to `static/arthur/` — an asset the project owns, not an extraction that churns. The server verifies rather than trusting a name it was handed: it recomputes every key from the descriptor stored beside it, refuses an analysis that declares no detector version, and refuses a document naming blocks it does not hold. It hashes the descriptor TEXT, because JS prints an integral double as `1` and Python as `1.0`, and a scheme where both ends re-render the numbers disagrees on the first parameter that happens to be whole. Two loose ends from step 8 closed on the way. `pack` no longer takes a `(track, frame)` predicate whose call sites each re-derived a feature from an index — every track names the feature it follows, which deleted five hand-maintained mappings. And `:dev-http` is gone: Django serves the page, shadow-cljs only builds into the staticfiles tree. 227 CLJS tests, 31 Django tests, green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 01:11:41 -04:00
```
In the app, upload a video, choose its footage, click **load frames**, then
**save**. Opening that project on another client reuses its saved landmarks and
mouth crops without detecting source frames again. The upload path derives its
footage response from database records; it does not create or consume a
`manifest.json` file. See [frontend/README.md](frontend/README.md) for details.
Anything under "## Run" and below describes the older JS prototype, which still
runs separately on port 8777.
### Deploy to Fly.io
The app uses a persistent Fly volume for its SQLite document database and
content-addressed footage blobs. Create the app once, set its Django secret, and
deploy from the repository root:
```sh
fly apps create arthur --org personal
fly volumes create data --region iad --size 1
fly secrets set DJANGO_SECRET_KEY="$(openssl rand -hex 32)" --app arthur
fly deploy --app arthur
```
The deployed app is at <https://arthur.fly.dev>. The container builds the
ClojureScript frontend, collects static assets, and runs database migrations on
startup.
Serve the document from a Django backend, split into three tiers Step 9. The tier split was the work; Django was the easy half. Tier 1 — the authored scene — is the document, and it is addressed as independently versioned leaves rather than saved whole, so one vertex drag cannot clobber a collaborator's keying. `domain/leaf` is the document as path -> value; `domain/wire` puts it on the wire as transit, because JSON has neither integer map keys nor keywords and a save would quietly turn `{0 v}` into `{"0" v}`. Tier 2 — the dense channel blocks — is content-addressed by a hash over every input, with the detector version inside every key through the analysis the block descriptor names. `flow/address`'s `block-knobs` is the invalidation table, and `address-test` does not trust it: it re-freezes the take once per knob and asserts the biconditional, that a block's bytes changed if and only if its key changed. That found `brow-pos` not depending on `contour-avg` — the brow ring is smoothed, the raise is not. Tier 3 — frames and audio — is served by the hash of its bytes out of the same store. A manifest now names frames and carries a URL for each, so the frame layout stopped being a shared secret between a shell script and a ClojureScript namespace, and the `?v=` cache-buster went with it: a blob's name is the hash of its contents, so a stale copy is not a thing that can happen. The synthetic take's `audio.wav` moved to `static/arthur/` — an asset the project owns, not an extraction that churns. The server verifies rather than trusting a name it was handed: it recomputes every key from the descriptor stored beside it, refuses an analysis that declares no detector version, and refuses a document naming blocks it does not hold. It hashes the descriptor TEXT, because JS prints an integral double as `1` and Python as `1.0`, and a scheme where both ends re-render the numbers disagrees on the first parameter that happens to be whole. Two loose ends from step 8 closed on the way. `pack` no longer takes a `(track, frame)` predicate whose call sites each re-derived a feature from an index — every track names the feature it follows, which deleted five hand-maintained mappings. And `:dev-http` is gone: Django serves the page, shadow-cljs only builds into the staticfiles tree. 227 CLJS tests, 31 Django tests, green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 01:11:41 -04:00
### How it is stored
Three tiers, cut by mutability and size — the full argument is in
[docs/architecture.md](docs/architecture.md):
| Tier | What | Where |
| --- | --- | --- |
| 1 **authored** | the scene: nodes, channels, features, time maps | the database, as independently addressed leaves. Kilobytes |
| 2 **derived** | detected landmarks, raw mouth crops, and dense channel blocks | `var/blobs`, addressed by analysis and block inputs, including the detector version |
| 3 **source** | the uploaded video, H.264 proxy and elementary stream, tracing stills, and audio | the same blob store, by the hash of their bytes |
Serve the document from a Django backend, split into three tiers Step 9. The tier split was the work; Django was the easy half. Tier 1 — the authored scene — is the document, and it is addressed as independently versioned leaves rather than saved whole, so one vertex drag cannot clobber a collaborator's keying. `domain/leaf` is the document as path -> value; `domain/wire` puts it on the wire as transit, because JSON has neither integer map keys nor keywords and a save would quietly turn `{0 v}` into `{"0" v}`. Tier 2 — the dense channel blocks — is content-addressed by a hash over every input, with the detector version inside every key through the analysis the block descriptor names. `flow/address`'s `block-knobs` is the invalidation table, and `address-test` does not trust it: it re-freezes the take once per knob and asserts the biconditional, that a block's bytes changed if and only if its key changed. That found `brow-pos` not depending on `contour-avg` — the brow ring is smoothed, the raise is not. Tier 3 — frames and audio — is served by the hash of its bytes out of the same store. A manifest now names frames and carries a URL for each, so the frame layout stopped being a shared secret between a shell script and a ClojureScript namespace, and the `?v=` cache-buster went with it: a blob's name is the hash of its contents, so a stale copy is not a thing that can happen. The synthetic take's `audio.wav` moved to `static/arthur/` — an asset the project owns, not an extraction that churns. The server verifies rather than trusting a name it was handed: it recomputes every key from the descriptor stored beside it, refuses an analysis that declares no detector version, and refuses a document naming blocks it does not hold. It hashes the descriptor TEXT, because JS prints an integral double as `1` and Python as `1.0`, and a scheme where both ends re-render the numbers disagrees on the first parameter that happens to be whole. Two loose ends from step 8 closed on the way. `pack` no longer takes a `(track, frame)` predicate whose call sites each re-derived a feature from an index — every track names the feature it follows, which deleted five hand-maintained mappings. And `:dev-http` is gone: Django serves the page, shadow-cljs only builds into the staticfiles tree. 227 CLJS tests, 31 Django tests, green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 01:11:41 -04:00
Only tier 1 is the document. Tier 2 is a pure function of tiers 1 and 3, so a
saved project names its blocks rather than carrying them, and a knob change gives
a block a new name rather than overwriting an old one.
## Run
```sh
Paint: vector background cels on the kept frames A sketch, and labelled as one. It exists to test whether the aesthetic holds when a human draws the background rather than the tracker deriving it, and it is meant to be replaced by a real paint surface with onion skin and undo. Kept to one dependency-free module so throwing it away is a delete, not surgery. Cels are drawn on the frames that get their own drawing and hold until the next one - the same rule the plate follows, and literally the same lookup, so the two cannot disagree about what is on screen. Editing always targets the cel you can see, so you can scrub anywhere and keep drawing. Pen places vertices and closes on the first one. Edit drags vertices or whole shapes, shift-click inserts, alt-click removes. Layers stack front-at-top with per-layer colour, show/hide and reorder. Drag a strip thumbnail onto the canvas to seed this cel from that one, every layer, as a deep copy - sharing the point arrays would make two cels silently edit each other. Two rules are enforced rather than left to discipline: Colours are PALETTE INDICES, never RGB. Sampling colour from the source is the one move docs/design.md calls irrecoverable, and a paint tool is exactly where that discipline would leak, so the picker cannot express a colour outside the ramp. Vertices snap to the 320x200 grid. On a hard-edged indexed rasteriser a shape nudged by 0.4px moves an edge by a whole pixel or not at all depending on where it happens to land, so sub-pixel vertices shimmer instead of holding still. Drawings autosave to localStorage per take name. They are the only thing here a person made by hand; everything else regenerates. Not in the .take export yet. Also adds serve.py, a no-store dev server. python3 -m http.server sends Last-Modified and browsers cache ES modules on it hard enough that a reload serves a stale app.js against a fresh index.html: the new knobs appear, nothing wires them, no error fires, and it reads as "your feature does not work". That cost real time this session. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 19:08:07 -04:00
python3 serve.py # from this directory, then open 127.0.0.1:8777
```
Paint: vector background cels on the kept frames A sketch, and labelled as one. It exists to test whether the aesthetic holds when a human draws the background rather than the tracker deriving it, and it is meant to be replaced by a real paint surface with onion skin and undo. Kept to one dependency-free module so throwing it away is a delete, not surgery. Cels are drawn on the frames that get their own drawing and hold until the next one - the same rule the plate follows, and literally the same lookup, so the two cannot disagree about what is on screen. Editing always targets the cel you can see, so you can scrub anywhere and keep drawing. Pen places vertices and closes on the first one. Edit drags vertices or whole shapes, shift-click inserts, alt-click removes. Layers stack front-at-top with per-layer colour, show/hide and reorder. Drag a strip thumbnail onto the canvas to seed this cel from that one, every layer, as a deep copy - sharing the point arrays would make two cels silently edit each other. Two rules are enforced rather than left to discipline: Colours are PALETTE INDICES, never RGB. Sampling colour from the source is the one move docs/design.md calls irrecoverable, and a paint tool is exactly where that discipline would leak, so the picker cannot express a colour outside the ramp. Vertices snap to the 320x200 grid. On a hard-edged indexed rasteriser a shape nudged by 0.4px moves an edge by a whole pixel or not at all depending on where it happens to land, so sub-pixel vertices shimmer instead of holding still. Drawings autosave to localStorage per take name. They are the only thing here a person made by hand; everything else regenerates. Not in the .take export yet. Also adds serve.py, a no-store dev server. python3 -m http.server sends Last-Modified and browsers cache ES modules on it hard enough that a reload serves a stale app.js against a fresh index.html: the new knobs appear, nothing wires them, no error fires, and it reads as "your feature does not work". That cost real time this session. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 19:08:07 -04:00
Use `serve.py`, not `python3 -m http.server`. The latter sends `Last-Modified`
and browsers cache ES modules on it hard enough that a reload serves a stale
`js/app.js` against a fresh `index.html` — new knobs appear in the markup, nothing
wires them, no error is raised, and the symptom reads as "the feature does not
work". `serve.py` is the same server with `no-store`.
2026-09-24 15:49:05 -04:00
Static files and ES modules — no build step, no dependencies beyond MediaPipe's
wasm, which is fetched from a CDN on first use.
**Synthetic take** needs no video and exercises everything below detection.
For real footage:
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
Upload it in the app. `./extract.sh` still writes the old PNG-sequence bundle and
`ingest_bundle` still registers it, but footage ingested that way has no decodable
stream and the loader will say so — the measured pixels come out of the video now.
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
MediaPipe's wasm and `face_landmarker.task` are both local; nothing in detection
touches the network.
Detection reads the H.264 elementary stream with WebCodecs, one coded frame at a
time. The proxy has no B-frames, so decode order is frame order. Each decoded
frame reaches MediaPipe in VIDEO running mode at its footage timestamp. The
decoder and detector advance together, with a pause between frames so progress
can paint. Saved analyses reuse their stored crop pixels and measure them with
the same pauses.
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
This is what replaced the PNG sequence, which was 112MB for 7.6 seconds and would
be 1.1GB at the 900-frame limit. The proxy is 6MB, and the landmarks barely
notice: detected off decoded H.264 rather than off the PNGs, they moved at most
0.0033 of frame width.
The server's footage manifest records the proxy's frame rate and frame count;
the page reads that rate because a guessed fps desynchronises audio from picture.
Choosing a lower picture rate happens after analysis.
**Picture fps** decides how often the finished roto gets a new pose. Analyze all
source frames, then sample those frozen poses at 12, 24 or the source rate while
keeping the original duration and audio. **Exposure** can hold a drawing across
more than one picture slot. The tracing editor chooses source frames for cel
references separately. Shared timing is the useful default for mouth, eyes,
teeth and plate so their changes read as one performance.
Eyes: lids, blinking, line of sight Three parts per eye, stacked the way the mouth is - dark lash ring, sclera inside it, iris inside that, square pupil in the iris. A blink then costs nothing: when the lid shuts the traced ring goes flat and the lash line collapses to a lens, which is a closed eye, drawn correctly, for free. Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a quantised position - and that is where the stylisation lives. Line of sight. Gaze is the iris centre relative to the midpoint of the eye's two corners, in units of corner distance. Both corners are in RIGID, so the origin and the scale are immune to the performance being measured; against the lid ring's centroid instead, every blink would drag the origin down and fake a glance at the floor on exactly the frames where the eye is most visible. Both eyes share one gaze - at this size the difference between the two measurements is noise, not vergence, and independent per-eye noise reads as wall-eyed immediately. Openness stays per-eye so a wink survives. Gaze is then quantised to a pixel grid with a dwell, which is not a stylisation imposed on the truth: real eyes move in saccades, and the smooth drift left in the measurement is tracker noise plus head-compensation error. Snapping to a grid removes the noise and recovers the saccade in one operation. The iris is placed in the frame of the already-smoothed, already-subsampled lid ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any even budget keeps them at 0 and n/2 - so it cannot drift relative to its own eye. Size is authored from the take mean, never remeasured per frame: a radius that breathes by a fraction of a pixel flickers a pixel on and off around the whole silhouette. iris anchor toggles steady/free/locked, because how much the eye wanders turns out to be an aesthetic choice and not only a correctness one. Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not have: blink hold. A blink is one frame at 12fps and a single frame of closed eye reads as a dropped frame, so once the eye shuts it stays shut long enough to be legible. Detection accuracy is not the problem; legibility is. The pupil is a square because at three pixels a circle is a plus sign with the corners gnawed off, and it changes shape as it moves. Drawn from a rounded centre shared with the iris so it is exactly its nominal size on every frame. Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro would: the lid crops the iris at extreme gaze for free, so nothing has to clamp the gaze, which would flatten the performance at the extremes that carry it. Which iris block belongs to which eye is RESOLVED from geometry, not declared. A swap looks almost right - each eye still has a disc roughly where it belongs - so it survives an eyeball and then reads as a subtly wall-eyed character forever. Voted across every frame; the test feeds a deliberately swapped track. Also: exposure. Aesthetic sparseness was set by the extraction rate, which made the timing a property of a directory of PNGs - auditioning 12 against 24 meant re-ripping and re-detecting the whole clip. It is now a render-time grid, on 1s/2s/3s/4s, so the dense track keeps everything and the audio clock is untouched. The take format already carried an exposure field; it was never driven. Everything rides the same grid, because a head cutting on the odd frames while the mouth cuts on the even ones reads as two performances laid over each other. 41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a swapped track, a blink does not fake a change of gaze, a stencilled disc cannot spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure never reads a pose from the future. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
**Audio is the playback clock**: `frame = floor(audio.currentTime * fps)`. A slow
render loop therefore drops frames instead of drifting, and ½x / ¼x work by
setting `playbackRate` with the picture following for free.
## Shooting for it
Near-frontal, good light, consistent scale, head reasonably still. Hold a neutral
closed mouth for a second at the top of the take: that frame is picked
automatically as the neutral and drives calibration and the placeholder plate.
Out-of-plane head rotation cannot be stabilised away by a 2D similarity
transform — the `residual` readout rises when it happens. The pipeline's answer
is hand-drawn head plates, which this tool does not yet do.
## Knobs
| Knob | What it does |
| --- | --- |
| vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. |
Eyes: lids, blinking, line of sight Three parts per eye, stacked the way the mouth is - dark lash ring, sclera inside it, iris inside that, square pupil in the iris. A blink then costs nothing: when the lid shuts the traced ring goes flat and the lash line collapses to a lens, which is a closed eye, drawn correctly, for free. Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a quantised position - and that is where the stylisation lives. Line of sight. Gaze is the iris centre relative to the midpoint of the eye's two corners, in units of corner distance. Both corners are in RIGID, so the origin and the scale are immune to the performance being measured; against the lid ring's centroid instead, every blink would drag the origin down and fake a glance at the floor on exactly the frames where the eye is most visible. Both eyes share one gaze - at this size the difference between the two measurements is noise, not vergence, and independent per-eye noise reads as wall-eyed immediately. Openness stays per-eye so a wink survives. Gaze is then quantised to a pixel grid with a dwell, which is not a stylisation imposed on the truth: real eyes move in saccades, and the smooth drift left in the measurement is tracker noise plus head-compensation error. Snapping to a grid removes the noise and recovers the saccade in one operation. The iris is placed in the frame of the already-smoothed, already-subsampled lid ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any even budget keeps them at 0 and n/2 - so it cannot drift relative to its own eye. Size is authored from the take mean, never remeasured per frame: a radius that breathes by a fraction of a pixel flickers a pixel on and off around the whole silhouette. iris anchor toggles steady/free/locked, because how much the eye wanders turns out to be an aesthetic choice and not only a correctness one. Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not have: blink hold. A blink is one frame at 12fps and a single frame of closed eye reads as a dropped frame, so once the eye shuts it stays shut long enough to be legible. Detection accuracy is not the problem; legibility is. The pupil is a square because at three pixels a circle is a plus sign with the corners gnawed off, and it changes shape as it moves. Drawn from a rounded centre shared with the iris so it is exactly its nominal size on every frame. Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro would: the lid crops the iris at extreme gaze for free, so nothing has to clamp the gaze, which would flatten the performance at the extremes that carry it. Which iris block belongs to which eye is RESOLVED from geometry, not declared. A swap looks almost right - each eye still has a disc roughly where it belongs - so it survives an eyeball and then reads as a subtly wall-eyed character forever. Voted across every frame; the test feeds a deliberately swapped track. Also: exposure. Aesthetic sparseness was set by the extraction rate, which made the timing a property of a directory of PNGs - auditioning 12 against 24 meant re-ripping and re-detecting the whole clip. It is now a render-time grid, on 1s/2s/3s/4s, so the dense track keeps everything and the audio clock is untouched. The take format already carried an exposure field; it was never driven. Everything rides the same grid, because a head cutting on the odd frames while the mouth cuts on the even ones reads as two performances laid over each other. 41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a swapped track, a blink does not fake a change of gaze, a stencilled disc cannot spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure never reads a pose from the future. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
| exposure | How often the picture changes: on 1s, 2s, 3s, 4s. Rip dense, choose timing here. |
| mouth lead ±f | Shifts the performance tracks earlier against the audio and the head. `[` `]`. |
| contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. |
| anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform*. |
| closed-mouth cut | Aperture below which the mouth interior is emitted as `hidden`. |
| suggest tolerance | Max head movement before a new plate drawing is required. Affects **Suggest** only. |
## Teeth
MediaPipe has no landmarks inside the lips: the inner ring bounds the cavity and
everything within it is just pixels. So teeth come from the image.
The hazard is vertex correspondence — a traced contour reorders between frames
and boils. The way out for a blob specifically is **radial sampling**: march
outward from the blob's centroid along N fixed directions and take the last pixel
inside. Vertex *k* is then always "the extent in direction *k*", so correspondence
holds by construction, the vertex count is fixed, and temporal smoothing is well
defined with no reordering possible. It also produces a star-shaped reduction,
which is what flat blocks of colour want.
The **teeth measurement** panel shows exactly what is sampled: region dimmed,
kept pixels green, extracted contour amber. Tune against that, not the numbers.
| Knob | What it does |
| --- | --- |
| teeth contrast | Gate on the separation between the cavity's dark and bright class means. Otsu always returns *some* threshold, so this is what stops it inventing teeth in a dark mouth. |
| cavity erode | Pulls the sampled region in from the lip edge — MediaPipe's inner lip landmarks sit slightly outside the real opening, and lips are bright. |
| blob grow/erode | Resizes the found blob. An open pass always runs first to despeckle. |
| tongue reject | Drops pixels red relative to their own brightness. Teeth are near-neutral; tongue is not. |
| prefer upper | Biases component choice toward the top of the cavity. Area alone picks the tongue when the mouth is wide. |
| teeth vertices | Radial sample count. |
| teeth avg ±f | Temporal average over the contour. |
| teeth dwell | Frames a presence change must persist before it takes effect. |
**Tongue** as its own part would work the same way, gated on redness instead of
brightness and biased low rather than high. Not implemented: it is not visible in
the test footage, which reads as a dark cavity with a bright upper-teeth band.
Eyes: lids, blinking, line of sight Three parts per eye, stacked the way the mouth is - dark lash ring, sclera inside it, iris inside that, square pupil in the iris. A blink then costs nothing: when the lid shuts the traced ring goes flat and the lash line collapses to a lens, which is a closed eye, drawn correctly, for free. Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a quantised position - and that is where the stylisation lives. Line of sight. Gaze is the iris centre relative to the midpoint of the eye's two corners, in units of corner distance. Both corners are in RIGID, so the origin and the scale are immune to the performance being measured; against the lid ring's centroid instead, every blink would drag the origin down and fake a glance at the floor on exactly the frames where the eye is most visible. Both eyes share one gaze - at this size the difference between the two measurements is noise, not vergence, and independent per-eye noise reads as wall-eyed immediately. Openness stays per-eye so a wink survives. Gaze is then quantised to a pixel grid with a dwell, which is not a stylisation imposed on the truth: real eyes move in saccades, and the smooth drift left in the measurement is tracker noise plus head-compensation error. Snapping to a grid removes the noise and recovers the saccade in one operation. The iris is placed in the frame of the already-smoothed, already-subsampled lid ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any even budget keeps them at 0 and n/2 - so it cannot drift relative to its own eye. Size is authored from the take mean, never remeasured per frame: a radius that breathes by a fraction of a pixel flickers a pixel on and off around the whole silhouette. iris anchor toggles steady/free/locked, because how much the eye wanders turns out to be an aesthetic choice and not only a correctness one. Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not have: blink hold. A blink is one frame at 12fps and a single frame of closed eye reads as a dropped frame, so once the eye shuts it stays shut long enough to be legible. Detection accuracy is not the problem; legibility is. The pupil is a square because at three pixels a circle is a plus sign with the corners gnawed off, and it changes shape as it moves. Drawn from a rounded centre shared with the iris so it is exactly its nominal size on every frame. Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro would: the lid crops the iris at extreme gaze for free, so nothing has to clamp the gaze, which would flatten the performance at the extremes that carry it. Which iris block belongs to which eye is RESOLVED from geometry, not declared. A swap looks almost right - each eye still has a disc roughly where it belongs - so it survives an eyeball and then reads as a subtly wall-eyed character forever. Voted across every frame; the test feeds a deliberately swapped track. Also: exposure. Aesthetic sparseness was set by the extraction rate, which made the timing a property of a directory of PNGs - auditioning 12 against 24 meant re-ripping and re-detecting the whole clip. It is now a render-time grid, on 1s/2s/3s/4s, so the dense track keeps everything and the audio clock is untouched. The take format already carried an exposure field; it was never driven. Everything rides the same grid, because a head cutting on the odd frames while the mouth cuts on the even ones reads as two performances laid over each other. 41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a swapped track, a blink does not fake a change of gaze, a stencilled disc cannot spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure never reads a pose from the future. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
## Eyes
Three parts per eye, stacked the way the mouth is: a dark **lash ring**, the
**sclera** inside it, and the **iris** inside that, with a square **pupil** in
the iris. The dark ring outside a pale interior is what makes a flat shape read
as an opening rather than a blob, and it is why a blink costs nothing — when the
lid shuts, the traced ring goes flat and the lash line collapses to a lens,
which is a closed eye, drawn correctly, for free.
The **lids are a feature**, rotoscoped like the mouth: head-local, a key on every
frame, the same `contour avg` knob. They track the face, because the face is what
they are attached to.
The **iris is a primitive** — a disc at a quantised position — and that is where
the stylisation is.
### Line of sight
Gaze is the iris centre relative to the **midpoint of the eye's two corners**,
in units of corner distance. Both corners are in `RIGID`, which is the point:
the origin and the scale are built only from landmarks that do not move under
performance. Measure against the lid ring's centroid instead and every blink
drags that centroid down and fakes a glance at the floor, on exactly the frames
where the eye is most conspicuous.
**Both eyes share one gaze.** At 320×200 an iris is a handful of pixels and its
centre comes from five landmarks on an eye twenty pixels wide, so the difference
between the two measurements is noise, not vergence — and independent per-eye
noise reads as wall-eyed immediately, which is the most expensive artefact on a
face. Openness stays per-eye, so a wink survives.
Then the gaze is **quantised to a pixel grid with a dwell**, which is not a
stylisation imposed on the truth: real eyes move in saccades, holding a fixation
and then jumping. The smooth drift left in the measurement is tracker noise plus
head-compensation error, so snapping to a grid and requiring a dwell removes the
noise and recovers the saccade in one operation. The readout reports how many
distinct cells the iris ever occupies — three or four is a character who looks at
things, forty is an unquantised iris sliding around.
The iris is drawn at the socket read back off the **already-smoothed, already-
subsampled lid ring** — slots 0 and 8 of a 16-slot ring are the corners, and
subsampling to any even budget keeps them at output indices 0 and `n/2`. So the
iris is placed in the frame of the exact polygon it sits inside and cannot drift
relative to its own eye. Size is authored from the take's mean eye width, not
remeasured per frame: a radius that breathes by a fraction of a pixel flickers a
pixel on and off around the whole silhouette.
The iris is **stencilled to the sclera** and the pupil to the iris — the indexed
buffer is its own clip mask, the way Animator Pro would do it. So the lid crops
the iris at extreme gaze automatically, and nothing needs to clamp the gaze,
which would flatten the performance at exactly the extremes that carry it.
### Blinking
Openness is the lid gap over the corner distance — normalised, so one threshold
carries across takes and faces. It gets hysteresis and a dwell like the teeth,
plus one knob the teeth do not have: **blink hold**. A blink is 100–150ms, which
is one frame at 12fps, and a single frame of closed eye reads as a dropped frame
rather than as a blink. Animators draw a blink over two or three drawings for
that reason, so once the eye shuts it stays shut for `hold` frames.
### The pupil is a square
At this size a pupil is three pixels across, and a circle of radius 1.5 is not a
circle — it is a plus sign with the corners gnawed off, and it changes shape as
it moves. A square that size is a deliberate mark that stays the same mark
wherever it lands. It is drawn from a rounded centre shared with the iris, so it
is exactly its nominal size on every frame instead of spilling to the next pixel
on some and not others.
### Which iris is which
The refined mesh appends ten iris points, five per eye, and MediaPipe's own
left/right naming is viewer-relative in some places and subject-relative in
others. Getting it backwards swaps the irises, which looks *almost* right — each
eye still has a disc roughly where it belongs — so it survives an eyeball and
then reads as a subtly wall-eyed character forever. The pairing is therefore
**resolved from the geometry**, by voting each block's distance to each eye's
corner midpoint across every frame, and the selftest feeds it a track built the
other way round to prove it actually looks.
| Knob | What it does |
| --- | --- |
| eye vertices | Lid ring vertex budget, off a 16-slot ring. |
| lash line | How far the dark ring sits outside the lid, in pixels. |
| blink cut | Openness below which the eye is shut. Normalised by corner distance. |
| blink hold | Minimum frames a blink stays on screen. A one-frame blink is a dropout. |
| blink dwell | Frames a change must persist. Usually 0 — unlike the teeth, a real blink *is* one frame. |
| gaze gain | Exaggerates or damps the throw. Measured excursion is small; a character usually wants more. |
| gaze step | The pixel grid the iris snaps to. 0 = off, and then dwell does nothing either. |
| gaze dwell | How long a new cell must hold before it takes. Together with step, this is what makes saccades. |
| iris size | Diameter as a percentage of eye width. |
| pupil | Square pupil in whole pixels. 0 = off. |
Brows: traced ring, quantised raise A brow at 320x200 is fourteen pixels wide and three tall. Its shape carries almost nothing at that size; its height above the eye carries the expression, and a brow raise is the most legible beat on a face. So the ring is traced and the height is quantised - the split the eyes already got, where the lid is a traced feature and the iris a quantised primitive. The decomposition is the point. The traced ring already contains the real height, so adding a quantised raise on top would move the brow twice. The height is measured OUT of the ring, quantised, and put back, so the shape that renders is his at a height that snaps between a few levels and holds. Measured at both ends rather than as one number, because raise and tilt are different expressions out of one mechanism: both ends up is surprise, inner up alone is worry, inner down is anger. They share a dwell - the gaze quantiser, renamed quantizeSnap now that it has two callers - so the brow hits its pose in one frame instead of crawling into it with one end arriving before the other. Measured against the eye's corner midpoint, never its lid. Same trap the gaze origin has and worth avoiding twice: brows and lids move together constantly, so a brow that jumped on every blink would read as a tic. Rest pose from the take median rather than the neutral frame, for the reason gaze learned the hard way - that frame is picked by minimum mouth aperture and says nothing about the brows. Two correspondences resolved from geometry, not declared: which ring is which brow, and which end is the outer one. The second matters more - backwards, the tilt mirrors and worry renders as its own opposite, which reads as a directed performance choice rather than a bug and would never be questioned. Which EDGE is upper is deliberately left unresolved: it traverses the same ring the other way, an even-odd fill has no winding, and both ends still land on fixed slots. Also fixes a bug from the exposure work: the live render applied exposure to the plate and the mouth but not to the eyes, so on 2s the preview and the export disagreed. A preview that disagrees with the export is the one bug this tool cannot afford. perfIndex now exists as a named thing so the two paths cannot drift apart again. 91 -> 105 assertions. Ground truth on all four synthetic brow poses, tilt separating worry from anger by sign, a blink not faking a raise, and a shared dwell never emitting a half-raised brow. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:11:25 -04:00
## Brows
A brow at 320×200 is about fourteen pixels wide and three tall. Its **shape**
carries almost nothing at that size; its **height above the eye** carries the
expression, and a brow raise is the most legible beat on a face. So the ring is
traced and the height is quantised — the same split the eyes got, where the lid
is a traced feature and the iris a quantised primitive.
The decomposition matters. The traced ring already contains the real height, so
adding a quantised raise on top would move the brow twice. Instead the height is
measured *out* of the ring, quantised, and put back: the shape that renders is
his, at a height that snaps between a few levels and holds.
Height is measured at **both ends**, not as one number, because raise and tilt
are different expressions out of one mechanism — both ends up is surprise, inner
up alone is worry, inner down is anger. They share a dwell, so the brow hits its
pose in one frame instead of crawling into it with one end arriving first.
It is measured against the eye's **corner midpoint**, never its lid — the same
trap the gaze origin has, and worth avoiding twice: brows and lids move together
constantly, so a brow that jumped on every blink would read as a tic. The rest
pose comes from the take **median**, not the neutral frame, for the same reason
gaze does: that frame is chosen by minimum mouth aperture and says nothing
whatever about the brows.
Two correspondences are resolved from geometry rather than declared: which ring
is which brow, and which end of a ring is the outer one. The second matters more
— get it backwards and the tilt mirrors, so worry renders as its own opposite,
which reads as a directed performance choice and would never be questioned.
Which *edge* of the brow is the upper one is deliberately left unresolved: it
traverses the same ring the other way round, an even-odd fill has no winding,
and the two ends still land on fixed slots either way.
| Knob | What it does |
| --- | --- |
| brow vertices | Ring vertex budget, off a 10-slot ring. |
| brow weight | Thickens the ring outward. It needs it at three pixels tall. |
| brow raise gain | Exaggerates or damps the raise. |
| brow step | The pixel grid the height snaps to. 0 = off. |
| brow dwell | How long a new height must hold. Shared across both ends. |
## The plate is reference, not art
The plate layer has several representations because its job changes. Cycle with
<kbd>B</kbd>:
| Mode | For |
| --- | --- |
| photo dim / photo | **Drawing over.** The source frame, stabilised. |
| posterized | The footage quantised into the ramp — a look test asked of the source rather than of a drawing. |
| oval | A flat stand-in, to judge the mouth against something. |
| oval + photo | Checking the stand-in against the real head. |
| none | Mouth alone. |
Photo modes are **registered**: the frame is mapped into raster space through the
same transform chain the contours go through, so the head sits still and a
drawing traced from it is already aligned to the mouth. An unregistered underlay
would be decorative.
`MediaPipe`'s face oval is the *face* boundary — it cuts at the hairline and
excludes hair, ears, jaw underside and neck — so as a head silhouette it is an
egg by construction, and no landmark precision fixes that. Hence the photo.
**Save frame 4x** writes the current registered composite as a 1280×800 PNG to
draw on.
Paint: vector background cels on the kept frames A sketch, and labelled as one. It exists to test whether the aesthetic holds when a human draws the background rather than the tracker deriving it, and it is meant to be replaced by a real paint surface with onion skin and undo. Kept to one dependency-free module so throwing it away is a delete, not surgery. Cels are drawn on the frames that get their own drawing and hold until the next one - the same rule the plate follows, and literally the same lookup, so the two cannot disagree about what is on screen. Editing always targets the cel you can see, so you can scrub anywhere and keep drawing. Pen places vertices and closes on the first one. Edit drags vertices or whole shapes, shift-click inserts, alt-click removes. Layers stack front-at-top with per-layer colour, show/hide and reorder. Drag a strip thumbnail onto the canvas to seed this cel from that one, every layer, as a deep copy - sharing the point arrays would make two cels silently edit each other. Two rules are enforced rather than left to discipline: Colours are PALETTE INDICES, never RGB. Sampling colour from the source is the one move docs/design.md calls irrecoverable, and a paint tool is exactly where that discipline would leak, so the picker cannot express a colour outside the ramp. Vertices snap to the 320x200 grid. On a hard-edged indexed rasteriser a shape nudged by 0.4px moves an edge by a whole pixel or not at all depending on where it happens to land, so sub-pixel vertices shimmer instead of holding still. Drawings autosave to localStorage per take name. They are the only thing here a person made by hand; everything else regenerates. Not in the .take export yet. Also adds serve.py, a no-store dev server. python3 -m http.server sends Last-Modified and browsers cache ES modules on it hard enough that a reload serves a stale app.js against a fresh index.html: the new knobs appear, nothing wires them, no error fires, and it reads as "your feature does not work". That cost real time this session. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 19:08:07 -04:00
## Paint — background cels
**A sketch.** It exists to test whether the aesthetic holds when a human draws
the background instead of the tracker deriving it, and it is meant to be
replaced by a real paint surface with onion skin and undo. It is one dependency-
free module, `js/paint.js`, so throwing it away is a delete rather than surgery.
Cels are drawn on the frames that get their own drawing and **hold until the
next one** — the same rule the plate follows, and literally the same lookup. You
can scrub anywhere and keep drawing on the cel you can see; the header says
which one you are editing and how far it holds.
- **pen** — click to place vertices, click the green box on the first one (or
<kbd>Enter</kbd> / double-click) to close.
- **edit** — click a shape to select, drag a vertex or the whole shape,
<kbd>Shift</kbd>-click an edge to insert a vertex, <kbd>Alt</kbd>-click one to
remove it, <kbd>Del</kbd> to delete the layer.
- **Layers** stack Photoshop-style, front at the top, with per-layer colour,
show/hide and reorder.
- **Copy previous** brings the last drawing forward onto this frame. It means
the nearest earlier enabled frame that *actually has* a drawing, skipping the
empty ones — every frame is enabled until you thin the strip out, so the naive
rule resolved to `f-1` and it looked like it only ever copied the frame to the
left. **Drag a frame** from the strip onto the canvas to seed from any other
frame instead. Both deep-copy; the two cels never share point arrays.
- Frames carrying a drawing are marked **▣** in the strip, so you can see the
rhythm rather than having to remember it.
Paint: vector background cels on the kept frames A sketch, and labelled as one. It exists to test whether the aesthetic holds when a human draws the background rather than the tracker deriving it, and it is meant to be replaced by a real paint surface with onion skin and undo. Kept to one dependency-free module so throwing it away is a delete, not surgery. Cels are drawn on the frames that get their own drawing and hold until the next one - the same rule the plate follows, and literally the same lookup, so the two cannot disagree about what is on screen. Editing always targets the cel you can see, so you can scrub anywhere and keep drawing. Pen places vertices and closes on the first one. Edit drags vertices or whole shapes, shift-click inserts, alt-click removes. Layers stack front-at-top with per-layer colour, show/hide and reorder. Drag a strip thumbnail onto the canvas to seed this cel from that one, every layer, as a deep copy - sharing the point arrays would make two cels silently edit each other. Two rules are enforced rather than left to discipline: Colours are PALETTE INDICES, never RGB. Sampling colour from the source is the one move docs/design.md calls irrecoverable, and a paint tool is exactly where that discipline would leak, so the picker cannot express a colour outside the ramp. Vertices snap to the 320x200 grid. On a hard-edged indexed rasteriser a shape nudged by 0.4px moves an edge by a whole pixel or not at all depending on where it happens to land, so sub-pixel vertices shimmer instead of holding still. Drawings autosave to localStorage per take name. They are the only thing here a person made by hand; everything else regenerates. Not in the .take export yet. Also adds serve.py, a no-store dev server. python3 -m http.server sends Last-Modified and browsers cache ES modules on it hard enough that a reload serves a stale app.js against a fresh index.html: the new knobs appear, nothing wires them, no error fires, and it reads as "your feature does not work". That cost real time this session. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 19:08:07 -04:00
Two rules are enforced rather than left to discipline. Colours are **palette
indices**, so you cannot pick one that is not in the ramp — sampling colour from
the source is the one move `docs/design.md` says is irrecoverable. And vertices
**snap to the 320×200 grid**, because on a hard-edged indexed rasteriser a shape
nudged by 0.4px moves an edge by a whole pixel or not at all depending on where
it lands, which shimmers instead of holding.
Drawings autosave to `localStorage` per take name. They are the only thing in
the tool a person made by hand; everything else regenerates. They are not in the
`.take` export yet.
## Two kinds of sparseness
Sparseness has two unrelated causes, and conflating them was the original design
error here. **Aesthetic** sparseness is chosen from the full analyzed source
track at rendering time. **Labour** sparseness is a human drawing each cel and
selecting which source frames to use as tracing references.
Aesthetic sparseness is the **picture fps** control, with exposure available for
longer holds. Both happen after analysis, so auditioning 12 against 24 needs no
re-extraction or re-detection.
Eyes: lids, blinking, line of sight Three parts per eye, stacked the way the mouth is - dark lash ring, sclera inside it, iris inside that, square pupil in the iris. A blink then costs nothing: when the lid shuts the traced ring goes flat and the lash line collapses to a lens, which is a closed eye, drawn correctly, for free. Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a quantised position - and that is where the stylisation lives. Line of sight. Gaze is the iris centre relative to the midpoint of the eye's two corners, in units of corner distance. Both corners are in RIGID, so the origin and the scale are immune to the performance being measured; against the lid ring's centroid instead, every blink would drag the origin down and fake a glance at the floor on exactly the frames where the eye is most visible. Both eyes share one gaze - at this size the difference between the two measurements is noise, not vergence, and independent per-eye noise reads as wall-eyed immediately. Openness stays per-eye so a wink survives. Gaze is then quantised to a pixel grid with a dwell, which is not a stylisation imposed on the truth: real eyes move in saccades, and the smooth drift left in the measurement is tracker noise plus head-compensation error. Snapping to a grid removes the noise and recovers the saccade in one operation. The iris is placed in the frame of the already-smoothed, already-subsampled lid ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any even budget keeps them at 0 and n/2 - so it cannot drift relative to its own eye. Size is authored from the take mean, never remeasured per frame: a radius that breathes by a fraction of a pixel flickers a pixel on and off around the whole silhouette. iris anchor toggles steady/free/locked, because how much the eye wanders turns out to be an aesthetic choice and not only a correctness one. Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not have: blink hold. A blink is one frame at 12fps and a single frame of closed eye reads as a dropped frame, so once the eye shuts it stays shut long enough to be legible. Detection accuracy is not the problem; legibility is. The pupil is a square because at three pixels a circle is a plus sign with the corners gnawed off, and it changes shape as it moves. Drawn from a rounded centre shared with the iris so it is exactly its nominal size on every frame. Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro would: the lid crops the iris at extreme gaze for free, so nothing has to clamp the gaze, which would flatten the performance at the extremes that carry it. Which iris block belongs to which eye is RESOLVED from geometry, not declared. A swap looks almost right - each eye still has a disc roughly where it belongs - so it survives an eyeball and then reads as a subtly wall-eyed character forever. Voted across every frame; the test feeds a deliberately swapped track. Also: exposure. Aesthetic sparseness was set by the extraction rate, which made the timing a property of a directory of PNGs - auditioning 12 against 24 meant re-ripping and re-detecting the whole clip. It is now a render-time grid, on 1s/2s/3s/4s, so the dense track keeps everything and the audio clock is untouched. The take format already carried an exposure field; it was never driven. Everything rides the same grid, because a head cutting on the odd frames while the mouth cuts on the even ones reads as two performances laid over each other. 41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a swapped track, a blink does not fake a change of gaze, a stencilled disc cannot spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure never reads a pose from the future. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
So the mouth keeps **every** frame: it is traced, and therefore free. In limited
animation lip sync is routinely the densest element, on 1s, while heads hold on
2s and 3s.
The frame strip is the editing surface for the other half: which frames need
their own plate drawing. Everything starts kept; delete what you don't want.
**Suggest** runs error-tolerance decimation over head pose as a starting point,
then you hand-correct.
**Mouth lead** is not a correction for a bug. A centred moving average has no
phase lag, so smoothing does not delay anything — but it blurs onsets, and the
visually salient moment of a mouth opening moves later even though the mean does
not. Animators also draw mouth shapes one or two frames ahead of the sound as
standard practice. Only the performance tracks shift; the head stays with the
audio, since it is the mouth that should anticipate. The lead is baked into the
exported take, so the renderer never needs to know about it.
`contour avg` is a deliberate, bounded exception to "never smooth the contour" in
`docs/design.md`. That rule held while keys were sparse, because sampling
at velocity minima rejected detector noise for free. With a key on every frame it
does not, so a radius shorter than the shortest articulation worth keeping is
justified — at 12fps, articulation spans 3–6 frames and detector noise is
per-frame, so ±1 separates them and ±3 starts eating speech.
## Tests
```sh
chromium --headless --virtual-time-budget=8000 --dump-dom \
http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+'
```
Brows: traced ring, quantised raise A brow at 320x200 is fourteen pixels wide and three tall. Its shape carries almost nothing at that size; its height above the eye carries the expression, and a brow raise is the most legible beat on a face. So the ring is traced and the height is quantised - the split the eyes already got, where the lid is a traced feature and the iris a quantised primitive. The decomposition is the point. The traced ring already contains the real height, so adding a quantised raise on top would move the brow twice. The height is measured OUT of the ring, quantised, and put back, so the shape that renders is his at a height that snaps between a few levels and holds. Measured at both ends rather than as one number, because raise and tilt are different expressions out of one mechanism: both ends up is surprise, inner up alone is worry, inner down is anger. They share a dwell - the gaze quantiser, renamed quantizeSnap now that it has two callers - so the brow hits its pose in one frame instead of crawling into it with one end arriving before the other. Measured against the eye's corner midpoint, never its lid. Same trap the gaze origin has and worth avoiding twice: brows and lids move together constantly, so a brow that jumped on every blink would read as a tic. Rest pose from the take median rather than the neutral frame, for the reason gaze learned the hard way - that frame is picked by minimum mouth aperture and says nothing about the brows. Two correspondences resolved from geometry, not declared: which ring is which brow, and which end is the outer one. The second matters more - backwards, the tilt mirrors and worry renders as its own opposite, which reads as a directed performance choice rather than a bug and would never be questioned. Which EDGE is upper is deliberately left unresolved: it traverses the same ring the other way, an even-odd fill has no winding, and both ends still land on fixed slots. Also fixes a bug from the exposure work: the live render applied exposure to the plate and the mouth but not to the eyes, so on 2s the preview and the export disagreed. A preview that disagrees with the export is the one bug this tool cannot afford. perfIndex now exists as a named thing so the two paths cannot drift apart again. 91 -> 105 assertions. Ground truth on all four synthetic brow poses, tilt separating worry from anger by sign, a blink not faking a raise, and a shared dwell never emitting a half-raised brow. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:11:25 -04:00
Or open `selftest.html`. 105 assertions over the stages below detection, plus a
wiring cross-check: every `el('id')` in `app.js` must exist in `index.html`. A
knob wired in one but not the other throws during wiring, which aborts the rest
of the module and leaves a blank page — a symptom that points nowhere near its
cause, and which has happened twice.
The ring-simplicity check is the load-bearing one. Because `hold` parts *cut*
between poses instead of interpolating, a ring whose vertex order is wrong
self-intersects and renders as blocks meeting at corners — and it is invisible at
odd `verts/2` and obvious at even, so it needs an assertion rather than an
eyeball.
## Not done yet
Brows: traced ring, quantised raise A brow at 320x200 is fourteen pixels wide and three tall. Its shape carries almost nothing at that size; its height above the eye carries the expression, and a brow raise is the most legible beat on a face. So the ring is traced and the height is quantised - the split the eyes already got, where the lid is a traced feature and the iris a quantised primitive. The decomposition is the point. The traced ring already contains the real height, so adding a quantised raise on top would move the brow twice. The height is measured OUT of the ring, quantised, and put back, so the shape that renders is his at a height that snaps between a few levels and holds. Measured at both ends rather than as one number, because raise and tilt are different expressions out of one mechanism: both ends up is surprise, inner up alone is worry, inner down is anger. They share a dwell - the gaze quantiser, renamed quantizeSnap now that it has two callers - so the brow hits its pose in one frame instead of crawling into it with one end arriving before the other. Measured against the eye's corner midpoint, never its lid. Same trap the gaze origin has and worth avoiding twice: brows and lids move together constantly, so a brow that jumped on every blink would read as a tic. Rest pose from the take median rather than the neutral frame, for the reason gaze learned the hard way - that frame is picked by minimum mouth aperture and says nothing about the brows. Two correspondences resolved from geometry, not declared: which ring is which brow, and which end is the outer one. The second matters more - backwards, the tilt mirrors and worry renders as its own opposite, which reads as a directed performance choice rather than a bug and would never be questioned. Which EDGE is upper is deliberately left unresolved: it traverses the same ring the other way, an even-odd fill has no winding, and both ends still land on fixed slots. Also fixes a bug from the exposure work: the live render applied exposure to the plate and the mouth but not to the eyes, so on 2s the preview and the export disagreed. A preview that disagrees with the export is the one bug this tool cannot afford. perfIndex now exists as a named thing so the two paths cannot drift apart again. 91 -> 105 assertions. Ground truth on all four synthetic brow poses, tilt separating worry from anger by sign, a blink not faking a raise, and a shared dwell never emitting a half-raised brow. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:11:25 -04:00
Hand-drawn head plates and per-plate mouth slots (the strip
decides *which frames need one*, but you cannot yet supply the drawing); real
performer→character calibration (currently identity, fitting the face oval to the
canvas); the override layer; anything on the Animator Pro side. The plate is a
face-oval polygon per kept frame — it exists so the mouth has a face to read
against, not to look good.
Serve the document from a Django backend, split into three tiers Step 9. The tier split was the work; Django was the easy half. Tier 1 — the authored scene — is the document, and it is addressed as independently versioned leaves rather than saved whole, so one vertex drag cannot clobber a collaborator's keying. `domain/leaf` is the document as path -> value; `domain/wire` puts it on the wire as transit, because JSON has neither integer map keys nor keywords and a save would quietly turn `{0 v}` into `{"0" v}`. Tier 2 — the dense channel blocks — is content-addressed by a hash over every input, with the detector version inside every key through the analysis the block descriptor names. `flow/address`'s `block-knobs` is the invalidation table, and `address-test` does not trust it: it re-freezes the take once per knob and asserts the biconditional, that a block's bytes changed if and only if its key changed. That found `brow-pos` not depending on `contour-avg` — the brow ring is smoothed, the raise is not. Tier 3 — frames and audio — is served by the hash of its bytes out of the same store. A manifest now names frames and carries a URL for each, so the frame layout stopped being a shared secret between a shell script and a ClojureScript namespace, and the `?v=` cache-buster went with it: a blob's name is the hash of its contents, so a stale copy is not a thing that can happen. The synthetic take's `audio.wav` moved to `static/arthur/` — an asset the project owns, not an extraction that churns. The server verifies rather than trusting a name it was handed: it recomputes every key from the descriptor stored beside it, refuses an analysis that declares no detector version, and refuses a document naming blocks it does not hold. It hashes the descriptor TEXT, because JS prints an integral double as `1` and Python as `1.0`, and a scheme where both ends re-render the numbers disagrees on the first parameter that happens to be whole. Two loose ends from step 8 closed on the way. `pack` no longer takes a `(track, frame)` predicate whose call sites each re-derived a feature from an index — every track names the feature it follows, which deleted five hand-maintained mappings. And `:dev-http` is gone: Django serves the page, shadow-cljs only builds into the staticfiles tree. 227 CLJS tests, 31 Django tests, green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 01:11:41 -04:00
In the port specifically: the parameter UI and scoped regeneration (the model is
built, the controls are not); automatic per-feature detection, so presence still
comes from the full-face mask plus a manifest annotation; multiplayer, for which
step 9 built the addressing and none of the socket; and in-browser extraction, so
`extract.sh` plus `manage.py ingest_bundle` is still how footage arrives.
Two smaller things that are known and undecided. `measure/brows` takes no
`presence` where `measure/eyes` does, so an occluded brow affects the freeze mask
but not brow measurement, and occluded landmarks still enter contour smoothing —
asymmetric with the eyes, and it is not settled which way is right. And `open`
takes the most recently updated project and shows its first clip: there is no
project browser, and the runtime store holds one clip at a time.