Steps 2 and 3 land together because the model revisions in the middle changed
code from both, and splitting them now would invent intermediate states that
never built.
domain/channel value-at across framed/keyed/dense, plus a cursor
domain/node decomposed transform, composition order, time maps
domain/scene topological order, z paths, eval-frame and resolver
clock audio-clocked frame derivation, outside app-db
db/events/subs re-frame arrives; the playhead is document state
ui/player the rAF loop; reads, blits, dispatches (almost) nothing
ui/shell transport
133 tests, 1158 assertions. The scene plays at 30fps against audio, scrubs, and
runs at 1/4x through 4x; verified by driving a real browser over CDP rather than
by assertion.
Two evaluators, on purpose. `eval-frame` is the specification -- allocating,
order-free, obviously correct. `resolver` is what playback uses: cached topo
order and z paths, a cursor per channel, a preallocated point buffer per node.
Both run the same walk, parameterised only by how a channel is read and where
points are written, because two independent implementations of frame evaluation
would drift and the drift would read as a rendering bug rather than as two
functions disagreeing. scene-test asserts they agree frame for frame in forward,
backward and random order.
Deviations and decisions, each with a reason:
- raster/fill-poly! is now a thin wrapper over fill-poly-buf!, which takes a flat
preallocated buffer. ONE scanline fill serves the analysis stages, which speak
{:x :y}, and frame evaluation, which hands over a buffer it owns. The parity
suite still passes pixel-for-pixel, which is what makes the rewrite safe.
- The state mask carries ABSENCE ONLY. An earlier draft gave it a hidden bit too,
per architecture.md's "hidden flag + palette index", and that bit was a dense
[:vis] wearing a different hat -- two mechanisms for one question, which is how
a part ends up hidden by one and shown by the other.
- The palette is a parameter of evaluation, not a global. A node names a TONE;
which ramp that tone is read in belongs to the timeline it sits in.
- :over layers and a symbol :rate THROW rather than being ignored. Neither is
built and nothing can produce one, so this can only fire on data that has run
ahead of the code. A silently dropped override is a hand correction the user
made once, watched fail, and has no reason to trust again.
Three findings the model produced rather than received:
- Presence propagates asymmetrically. An absent transform drops the subtree; an
absent [:geom :pts] drops only that node, because an absent mouth outline has
nothing to draw but the head it hangs off has not moved. That asymmetry is the
reason presence is tracked per channel and not per node.
- Z paths need lexicographic compare, not `compare`, which orders vectors by
count first -- so a cel three levels under "a1" would jump in front of a bare
"a2" and the layer order would mostly work.
- A node stencilled by something that drew nothing is dropped, not drawn
unclipped: an iris floating over the cheek is worse than a missing iris.
docs/ revised alongside, and those revisions are the load-bearing part:
- A scene, a timeline and a symbol are one type. The doc had two structures with
the same fields and never said so. Two axes of nesting are now separated --
parent/child within a timeline is flat with parent pointers, instance nesting
is by reference -- which is why "nestable" and "flat" only sounded
contradictory.
- Palettes are named, live on the project, and are ENABLED on a timeline as a
channel. Absent inherits; present travels with the timeline, so a symbol
authored against :night stays night wherever it is placed. The output index
space is the concatenation of the named ramps, which keeps one buffer and one
flat table and incidentally stops two nodes in different palettes colliding on
a stencil.
- Stabilisation is a channel, not a mode: {s, theta, tx, ty} IS [:xform :*], so
the normalise on/off/per-plate toggle is which of the three channel shapes the
:head node carries. Always measure and always store factored -- smoothing and
velocity-minimum key selection both need the split to exist in storage.
- There is no camera node and none is needed. Placement is a node transform, the
stage clips what hangs off it, and project dimensions are independent of the
footage. `makeXform` is therefore not to be ported: it bakes a cropping
decision into every stored vertex.
- Export is removed. The .take writer was for an Animator Pro render script; the
target is encoding video in the browser, and step 9 now says not to port the
old one.
demo/swarm is 120 shapes on six orbits, entirely dense blocks behind store
handles -- the shape freeze produces at step 5, and the first thing to exercise
that path under load. It plays at 30fps, and bench-test keeps a deliberately
loose floor under it because a performance regression here does not announce
itself: the picture stays correct and merely arrives late.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PDfHGdV39zu6rvgbBTfDaT
17 KiB
arthur — port plan and handoff
Self-contained. You should not need any prior conversation to execute this.
What arthur is
A tool that turns live-action video into 2D animation that reads as hand-authored: flat polygons, a tiny indexed palette, hard edges, 320×200, no antialiasing, motion carried by silhouette. It tracks a face out of a clip, reduces the lip contour to a handful of vertices, derives teeth from image content, and renders flat indexed fills.
It currently works, as vanilla JS ES modules with no build step. python3 serve.py, open 127.0.0.1:8777. Synthetic take exercises everything below
detection with no video needed.
This plan converts it to ClojureScript + re-frame, restructured around one uniform animation data model, and adds a Django backend for persistence and (later) collaboration.
Status of the existing documents
| File | What it is | Authority |
|---|---|---|
js/** |
the working tool, ~4,800 lines | authoritative. The comments encode bugs that actually happened. |
docs/animation-model.md |
the target data model: nodes, channels, symbols, time maps | build to this |
docs/architecture.md |
module layout, stages, sync and baking design | build to this; much of it is future scope |
docs/design.md, README.md |
prior synthesis by an earlier agent | useful, not authoritative. Revise freely. Do not treat its aesthetic claims as settled requirements. |
Where a document and the code disagree, the code wins, and the invariant list below is lifted from the code for exactly that reason.
Target repo layout
Both halves live here. Django at the root, because manage.py at the root is the
convention and keeps every python manage.py invocation working with no cd.
arthur/
mise.toml toolchain for both halves
manage.py
requirements.txt
server/ Django project: settings, urls, asgi, wsgi
clips/ Django app: models, views, consumers, routing, migrations
frontend/ the CLJS app
shadow-cljs.edn
package.json
src/arthur/** namespace root stays arthur.* whatever the dir is called
test/arthur/**
static/arthur/js/ shadow-cljs output, collected by Django staticfiles
docs/
js/ index.html serve.py extract.sh the old tool — see "the oracle"
clips is a naming call, not a constraint — it is the Django app holding
Project, Clip, Footage, Analysis, Leaf and Revision. Rename in one line if
something fits better.
Dev runs two processes: Django serves the page, shadow-cljs watch app rebuilds
into static/arthur/js. Set :output-dir "../static/arthur/js" in
shadow-cljs.edn.
Toolchain
mise install from the repo root. mise.toml pins java 21+, node 20, clojure,
python 3.12, and creates .venv.
Verified to resolve cleanly: reagent 1.2.0, re-frame 1.4.3, current
shadow-cljs.
Scope
In: analysis → keyframes → playback. The pure numeric core, the animation data model, a player, the measurement stages, and freezing measurements into channels.
Out, and do not build it: paint and cels; suggest (it only decides which
frames get a hand-drawn cel, so it has no job until drawing exists); the timeline
and sequences; symbols and the plate library; multiplayer; the override layer.
Each is designed for in docs/architecture.md and docs/animation-model.md.
Leave the :over field present and empty; leave :symbol out entirely.
The data model
Full specification in docs/animation-model.md. The subset to build:
;; The scene is a flat map of id -> node. Parent pointers, never nested maps.
{:id :mouth :kind :poly :parent :head :z "a3" :stencil nil :span [0 240]
:time {:mode :inherit} ; or {:mode :map :expose 2 :offset -1 :rate 1.0}
:channels
{[:xform :pos] {:animated? false :value [0.0 0.0]}
[:xform :rot] {:animated? false :value 0.0}
[:xform :scale] {:animated? false :value [1.0 1.0]}
[:xform :skew] {:animated? false :value [0.0 0.0]}
[:xform :anchor] {:animated? false :value [0.0 0.0]}
[:geom :pts] {:animated? true :interp :hold
:dense {:store "sha256:…" :offset 0 :stride 40 :frames 600}
:generated {:by :roto/lips-outer :analysis "sha256:…"
:params {:verts 8 :contour-avg 1}}
:over []}
[:style :color] {:animated? false :value :skin-dark}
[:vis] {:animated? false :value true}}}
Three channel shapes, one accessor (value-at channel f):
{:animated? false :value v}— static. A thing that simply exists.{:animated? true :interp :hold :keys {0 v, 4 v}}— sparse, authored, in the document. Keys are a map by frame, never a vector. Store a plain map (transit loses sortedness) and build the sorted index in the resolver.{:animated? true :interp :hold :dense {...}}— generated, one value per frame, in a typed array outside app-db.
:generated is provenance and the renderer never reads it. It is what the UI
uses to offer a parameter panel instead of raw keys. It lives on the channel,
not the node, because a node wants a rotoscoped [:geom :pts] and a
hand-animated [:xform :pos] at the same time.
:skew, :span, :anchor and :over stay in the shape even though nothing
drives them yet: each is a component of a decomposition or of a composition
order, and adding one later migrates every stored transform.
Transform composition, per node:
local = T(pos) · T(anchor) · R(rot) · K(skew) · S(scale) · T(-anchor)
world = world(parent) · local
What the prototype knows that you would otherwise rediscover
The JS is a prototype. Its conclusions about what looks right are provisional and you may revisit any of them; several contradict each other already. But a few things in it are not taste — they are facts about MediaPipe, about the maths, or about what an operation means — and those cost real time to rediscover.
Mechanical. Getting these wrong produces wrong output, not a different look.
- MediaPipe's normalised space is anisotropic. It divides x by image width
and y by height, so equal numbers do not mean equal pixels. Multiply x by
aspect = W/Hbefore any fit, or a "similarity" fitted in that space is not one and head roll comes out subtly wrong. When mapping pixels for an underlay, both axes divide byimgH. - MediaPipe's left/right naming is viewer-relative in some places and subject-relative in others. Any left/right pairing read off a table is a coin flip, and a swap looks almost right — each eye still has an iris roughly where it belongs — so it survives inspection. Resolve it from geometry.
- Ring tables are ordered traversals, and slot position is the vertex's
identity. That is what makes temporal correspondence possible at all, whatever
you decide the shapes should look like.
subsampleSlotsreturns ring positions, not landmark ids. - A wrongly-ordered ring self-intersects, and it is invisible at odd vertex budgets and obvious at even ones. If you keep ordered rings, assert simplicity in a test; no amount of looking will catch it reliably.
- Scaling a ring to thicken it collapses when the ring is degenerate — a shut eyelid scaled by 1.1 is still shut, so the lash line vanishes on exactly the frames where it is the whole drawing. A fixed radial offset does not. Maths, not taste.
- A fractional centre for a small integer-sized shape changes its size. Round the origin, not the extents, or a 3px mark is 3px on one frame and 4px on the next.
- Order of operations on time: flooring onto a grid and shifting against the clock do not commute. Shift first and the floor discards it on most frames.
Choices the prototype made. Revisit freely; here is what each was for.
| Choice | Its stated reason | How you would learn it was wrong |
|---|---|---|
| similarity (4 DOF), not affine | extra DOF absorbs out-of-plane head rotation as shear and smears it into the mouth | the residual readout stops responding to head turn |
| reference is the Procrustes mean over the shot, not frame 0 | no single frame's idiosyncrasies get baked into every other | one frame's detection error biases the whole take |
| smooth the transform, not the contour | sparse keys at velocity minima rejected detector noise for free | it was already broken by a "bounded exception" once keys went dense, so it was never a law |
| gaze measured against the eye's corner midpoint | measured against the lid, every blink drags the origin down and fakes a glance at the floor | gaze correlates with blinks |
| one gaze shared by both eyes | at this size the per-eye difference is noise, and independent noise reads as wall-eyed | a wink or a real vergence is lost |
| hold, never interpolate | a tweened mouth reads as puppet software | motion looks stepped rather than snappy |
| palette indices, never sampled RGB | sampling colour produces a pixel-art filter irrecoverably | — |
These are where to look first if the output is wrong. They are also where to look first if you want to change the look.
Conventions
domain/*may not requireflow/*; neither may requirere-frame.- Every flow function is
(f params inputs) -> output. No state, no db, no atoms. - Nothing below
subs/callssubscribe. - Every analysis function that reads pixels takes a
debug?flag and returns its intermediate masks alongside its result, the wayinterior.js/extractTeeth(..., wantDebug)already does. - Port the invariant comments across verbatim. They are the most valuable text in the repo.
The oracle
Keep js/, index.html and serve.py in the tree through step 5. They cost
nothing, serve.py still runs the old tool, and they are the numeric oracle:
run both implementations on the same synthetic track and diff.
fit-similarity and procrustes-mean should agree to 1e-9; a larger gap is a
port bug, not float noise.
Parity proves the port is faithful, not that the answer is right. The JS is a prototype, so keep the two kinds of test apart: a parity test pins behaviour while you move it, and is deleted once the move is done; a correctness test asserts something you have decided you want, and stays. Conflating them bakes the prototype's mistakes into the rewrite and makes them permanent. Delete them in one commit once the CLJS player renders the synthetic take correctly.
Do not port the debug views (drawPanes, drawInteriorDebug,
drawEyeOverlay, drawGazeDebug in js/app.js). The knowledge in them is not
the canvas calls — it is which things you must see to tune teeth: the source
crop, the in-region mask, the surviving mask, and the local contour. That contract
already exists as extractTeeth(..., wantDebug) returning
debugCanvas(src, inReg, mask, pw, ph, local). Port the payload, skip the
drawing. Redrawing it is ten lines whenever it is wanted.
Steps
Each step ends somewhere runnable. Do not proceed past a step whose "done" does not hold.
0 — scaffold and the oracle
mise install. Create frontend/ with shadow-cljs, reagent, re-frame. Port
synth.js (the synthetic landmark generator, including its swapIris flag) and
the numeric assertions from selftest.js to cljs.test.
Done: the suite runs and fails informatively.
1 — the pure bottom
Port verbatim: landmarks.js → domain/landmarks, mathutil.js → domain/geom,
ring helpers → domain/ring, raster.js → domain/raster, the palette →
domain/palette.
Done: tests pass, including ring simplicity and the swapped-iris vote. Numeric agreement with the JS to 1e-9. Nothing renders.
2 — the data model, with no analysis in it
domain/channel (value-at across all three shapes, plus a per-channel cursor),
domain/node (transform composition), domain/scene (topological order by parent
depth, eval-frame → draw ops in z order).
Hand-write a scene in EDN — a rectangle parented to a group whose
[:xform :pos] is keyed on four frames — and render it through domain/raster
into a canvas.
This is deliberately before any analysis. The data model has never been validated; find out here, with fifty lines to throw away, rather than after porting nine hundred lines of measurement into a shape that does not work.
Done: something moves on screen.
3 — the player
clock (audio-clocked: frame = ⌊currentTime · fps⌋, so a slow loop drops frames
instead of drifting; ½× and ¼× come free from playbackRate), the rAF loop, a
::resolver sub, and transport UI.
The loop reads and blits and dispatches nothing. The sub yields a resolver closure; the loop applies it at the playhead. The playhead itself lives in app-db like everything else — with layer-2 extractors and layer-3 computations, a playhead tick does not invalidate the expensive stages.
Done: the hand-written scene plays at 30fps against audio, scrubs, and runs at ½× and ¼×.
4 — measure: anchor and mouth
Port stabilize and the lip rings out of pipeline.js. Split condition
(smoothTransforms, smoothContours) into its own stage so the two smoothing
knobs do not re-run measurement.
Done: measured numbers match the JS on the synthetic track. Note that
parity here is on stabilize's output, not on toRasterRing's — the framing
step is being deleted, not ported.
5 — freeze
The new module, and the heart of this work: measurements → channels. A dense
[:geom :pts] block per node — Int16Array[frames × verts × 2] in the node's
own local space, with the block's fixed-point scale in its header — plus
:generated. Fixed topology is what makes this a rectangular array with no
per-frame header.
Do not port makeXform. The prototype bakes the framing into the stored
numbers: it centres on the face oval's bbox and zooms until the face is 80% of
the raster height, so every vertex carries a cropping decision made once from one
frame's landmarks. Dropping it is a deletion. Placement becomes [:xform :*] on
an authored :face node, the stage clips whatever hangs off, and project
dimensions stop being tied to the footage. See "What space geometry is in" in
docs/animation-model.md.
The anchor transform freezes onto :head, one level under :face, and the
normalise on/off/per-plate toggle is which of the three channel shapes that node
carries. Always measure and always store factored, whatever the toggle says:
smoothing and velocity-minimum key selection both require the split to exist in
storage.
Done: the synthetic take plays back as a moving mouth. Full vertical slice.
6 — detect
MediaPipe interop behind one namespace; real frames, real audio, real fps from the manifest. Vendor the wasm rather than fetching from jsdelivr — it is currently the only thing in the tool that silently requires a network.
Done: real footage plays back as a rotoscoped mouth.
7 — the rest of measure
Eyes (openness, gaze, iris pairing vote, blink resolution with its hold), brows
(raise and tilt at both ends, both correspondence votes), interior (otsu,
morphology, components, radial contour). Each keeps its debug? payload.
Two things fall out of the model instead of being written: the brow's
measure-the-height-out-and-put-it-back is [:geom :pts] plus [:xform :pos], two
channels on one node; and the iris is a :disc node parented to the lid ring and
stencilled by the sclera.
Done: parity with the JS tool, minus paint.
8 — knobs
The parameter UI, as leaf-addressed params in app-db
(clip/:cid/params/:subject/:area — see below), so the sync layer added later has
nothing to retrofit.
9 — backend
Django project, the clips app, models for Project/Clip/Footage/Analysis/Leaf,
and project load/save. Round-tripping a project through the server is the proof
the model serialises.
Output is not in this plan. The .take writer in js/take.js was for
driving an Animator Pro render script and it is not where this is going: the
target is encoding video in the browser, and that is a separate piece of design
nobody should pre-empt by porting the old one.
Two things to not foreclose
The feature controls will later be rethought to handle more than one face, periodic occlusion, stable identity across frames, and feature groups with their own parameters. That design can wait; two decisions here are free now and annoying to reverse:
- Presence is not visibility. An occluded subject has no value on a frame,
which is different from a part being hidden. Give every dense block a
Uint8Arraystate mask per frame and let it mean absent as well as hidden. - Params carry a subject segment.
clip/:cid/params/:subject/:area, with one subject today. Adding a path segment later touches every read and write.
The identity tracker, when it comes, should use the same pattern the iris and brow correspondences already use: vote across every frame rather than trusting one.