arthur/frontend/README.md
Olive Vaughn 35ef150b48 Give features stable identity, eye pairs and per-feature presence
Step 8's data model, ahead of its controls. Nothing here is a UI.

domain/params holds every knob's definition once — default, applicable area,
value constraints and the areas a change would force to regenerate. flow/take's
literal knob map becomes a view of it, so the take's defaults and the future
parameter panel cannot drift apart.

domain/feature adds subjects, features and groups as document data the renderer
never reads. A feature ID is stable for the whole clip, across occlusion: a run
of visible frames is not a new identity. An eye pair is an explicit group of one
or two eyes of the same subject, so a profile view with one identified eye needs
no invented partner. Settings resolve area -> subject -> group -> feature, and
dropping an eye from a pair materialises its effective values first so playback
does not jump. scene/problems now validates all of it.

Presence becomes per-feature rather than per-subject. freeze's :absent predicate
takes a track as well as a frame, so one occluded eye can be absent while its
partner still has a value; a full-face miss still marks everything absent. A
manifest may annotate known gaps as one-based inclusive intervals, which ingest
expands into observation tracks before measurement. An unobserved eye then gets
no vote in the iris pairing and cannot steer the shared gaze — gaze falls back to
whichever eye is visible. Temporal filters still see a sample on every frame,
held from the last observed one, because the numbers are a rectangular buffer;
the state mask, not the buffer, is what says the frame has no value.

js/app.js gets the same occlusion lesson: leading nulls from a face that starts
occluded used to throw away the whole take, and the neutral frame could be chosen
from a held duplicate pose.

Parameter editing, scoped regeneration and a feature-level detector remain. Until
one exists, footage without annotations falls back to the full-face mask rather
than claiming occlusions it cannot see.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B87NVmiU36qQmN9gmFYnJ9
2026-09-27 22:44:36 -04:00

187 lines
8.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# frontend
The ClojureScript half. See `docs/port-plan.md` for what is being built and in
what order; this file is only how to run it.
## Once
```sh
mise install # from the REPO ROOT: java 21+, node 20, clojure, python
cd frontend && npm install
```
`java` must be 21+. On an older JDK shadow-cljs fails with "CompilerOptions has
been compiled by a more recent version of the Java Runtime", which reads like a
shadow-cljs bug and is not one. `mise install` is what prevents it.
## The tests
```sh
cd frontend && mise exec -- npm test
```
Two things: compile the `:test` build, run it under node.
```
shadow-cljs compile test
node out/node-tests.js
```
Run them separately if a compile error is in the way.
**Run them through `mise`**, or make sure `mise`'s node is first on PATH. `java`
must be 21+ and node 20.19+. On an nvm node 20.11 shadowing the pinned one,
things fail in ways that read like the code being broken and are not.
### And the browser one
Step 5's done-criterion is a PICTURE, and no assertion in `cljs.test` can check
one: a take that resolves to the right numbers and draws nothing would pass every
test in `arthur.flow.freeze-test`. A blank canvas under a perfectly correct
transport is the bug class unit tests miss, and it has happened here once.
So there is a second suite that drives a real Chrome over CDP. It needs the dev
server up:
```sh
cd frontend && mise exec -- npx shadow-cljs watch app # in one shell
cd frontend && mise exec -- npm run browser # in another
```
No dependencies. Playwright is not installed and CDP needs none —
`node --experimental-websocket` has a global `WebSocket` and
`--headless=new --remote-debugging-port=N` is the whole of the other side. It
reads the canvas's own pixels rather than a screenshot, because the CSS scales
the stage up by 2 and a screenshot is four pixels per raster pixel; it writes
PNGs into `test/browser/out/` anyway, so "it drew something" can be checked by
eye as well as by count.
## The app
```sh
cd frontend && mise exec -- npx shadow-cljs watch app
```
Then open **<http://localhost:8778/index.html>** — with the `/index.html`, not
bare `/`. This shadow-cljs does no directory-index resolution, so `/` is a 404
whatever the roots are.
Four built-in clips, on buttons in the transport:
| | |
| --- | --- |
| `take` | the synthetic take, head **as filmed**. Step 5's deliverable: a moving mouth, frozen into dense channels, with no video file anywhere. |
| `locked` | the same freeze, head **locked**. The same blocks — `:head`'s channels are written as framed identity instead of as a dense track, and nothing in tier 2 differs. |
| `demo` | the hand-written scene from step 2. Not a face: the smallest scene that exercises every mechanism the model claims to have, so that each one is visible when it breaks. |
| `swarm` | a hundred and twenty dense nodes. Not useful; it is the load test. |
`take` and `locked` are the pair worth looking at together, because switching
between them is the whole of what "stabilisation is a channel, not a mode" means.
The demo scene itself is `src/arthur/demo/scene.edn`. Both the synthetic take
and real footage use `src/arthur/flow/take.cljs` for the measurement order and
`src/arthur/flow/freeze.cljs` for the landmark-to-channel conversion.
### Real footage (port steps 6–7)
From the repo root, extract a clip, then click **load frames** in the CLJS app:
```sh
./extract.sh /path/to/clip.mov
```
This keeps every source frame and writes `frames/0001.png` onward, `audio.wav`,
and `manifest.json` at the repo root. The manifest supplies the exact frame
count, source fps and audio path. Variable frame rate sources are rejected until
the manifest and clock carry per-frame timestamps.
To keep multiple takes or compare with a previous extraction, pass a bundle
directory and enter its manifest path in the app:
```sh
./extract.sh /path/to/clip.mov scratch/my-take
# source manifest field: /scratch/my-take/manifest.json
```
`scratch/` is ignored by Git. The directory contains its own frames, audio and
manifest, so extracting it does not replace another take's files.
Loading detects one face per frame, measures the mouth, eyes and brows from
landmarks and the teeth from source pixels, then freezes them into channels,
and adds a button for the footage clip. Detection happens once when you load;
playback only resolves channels and paints. Frames without a detection remain
marked absent even though their neighbouring poses are used to condition the
track. The scene now records stable subject and feature IDs and explicit eye
pairs; dense channels can mark one feature absent while another is observed.
Current MediaPipe loading supplies only the full-face detection mask. The stage
stays 320×200 regardless of the footage dimensions. Real
footage starts at the source picture rate. The **picture fps** buttons sample the
frozen roto at lower rates while the source track, duration and audio clock stay
unchanged. Picking frames to trace into cels is a separate future editing step.
For known occlusion intervals, an extracted manifest may add
`"feature-absence": {"eye-r": [[10, 14]]}`. Frame numbers are one-based and
inclusive, matching PNG filenames. The eye remains the same feature when it
reappears; the other eye and the mouth continue through the gap. This is an
input annotation, with no UI for editing it yet.
MediaPipe's JS, wasm and model are under `public/mediapipe/` and served locally.
No CDN is used by this app. See that directory's README for provenance.
Port 8778 is deliberately not 8777. `python3 serve.py` from the repo root still
runs the old JS tool on 8777, and the two are meant to run side by side.
From step 9 Django serves the page and `:dev-http` goes away.
## The oracle, which is finished
`js/` was the numeric oracle through step 4: `test/parity/` ran both
implementations on the same synthetic track and diffed `fit-similarity`,
`procrustes-mean`, the raster and `stabilize` to 1e-9.
**It was deleted at step 5, on purpose.** Parity proves the port is FAITHFUL, not
that the answer is RIGHT. The JS is a prototype and several of its conclusions
contradict each other; a parity test pins behaviour while code moves, and keeping
it afterwards would bake the prototype's mistakes into the rewrite and make them
permanent. `docs/port-plan.md` says to delete it in one commit once the CLJS
player renders the synthetic take, and that is what happened.
`js/` itself stays as the reference for the MediaPipe setup, face measurements
and pixel extraction. Its comments encode bugs that actually happened.
## Layout
```
src/arthur/domain/ pure. No re-frame, no DOM, no flow/.
src/arthur/flow/ the stages. `(f params inputs) -> output`, no state.
src/arthur/synth.cljs the synthetic track. In src/ because the take PLAYS it —
it stands in for flow/detect, and a tool that needs a
video file before it shows you anything is one you
cannot debug.
src/arthur/demo.cljs the hand-written scene, read from demo/scene.edn
src/arthur/demo/take.cljs the synthetic source for the shared flow/take path
src/arthur/ui/canvas.cljs indexed raster blit to the display canvas
test/browser/ drives a real Chrome over CDP. Not run by `npm test`.
public/index.html dev host page. Django replaces it at step 9.
```
## Two evaluators, on purpose
`domain/scene` has both `eval-frame` and `resolver`, and they are not
alternatives:
- **`(eval-frame scene f store)`** is the specification. Allocating, order-free,
obviously correct. Tests and one-off renders use it.
- **`(resolver scene store)` -> `(fn [f] ops)`** is what playback uses. It caches
the topological order and the z paths, holds a cursor per channel and reuses
one point buffer per node, so a frame allocates the op maps and nothing else.
Both run the same walk, parameterised by how a channel is read and where its
points are written — two independent implementations of frame evaluation would
drift, and the drift would look like a rendering bug rather than like two
functions disagreeing. What differs between them is exactly the part that can be
wrong, and `scene-test` asserts they agree frame for frame in forward, backward
and random order.
Because the resolver reuses its buffers, **ops must be rasterised before the
next frame is asked for.** That is the contract the rAF loop wants anyway: it
reads, blits, and dispatches nothing.