arthur/frontend/README.md
Olive Vaughn 83d106bbc5 Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.

Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:

/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.

A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.

Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.

Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.

The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.

Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00

14 KiB
Raw Blame History

frontend

The ClojureScript half. See docs/port-plan.md for what is being built and in what order; this file is only how to run it.

Once

mise install                                  # from the REPO ROOT
pip install -r requirements.txt               # the Django half; one dependency
mise exec -- python manage.py migrate         # the document database
cd frontend && npm install

java must be 21+. On an older JDK shadow-cljs fails with "CompilerOptions has been compiled by a more recent version of the Java Runtime", which reads like a shadow-cljs bug and is not one. mise install is what prevents it.

The tests

cd frontend && mise exec -- npm test

Two things: compile the :test build, run it under node.

shadow-cljs compile test
node out/node-tests.js

Run them separately if a compile error is in the way.

Run them through mise, or make sure mise's node is first on PATH. java must be 21+ and node 20.19+. On an nvm node 20.11 shadowing the pinned one, things fail in ways that read like the code being broken and are not.

And the Django one

mise exec -- python manage.py test clips      # from the REPO ROOT

Tests cover the API: the blob store, key verification, the load/save round trip, the conditional write, source analysis blocks, and video upload and extraction. The two groups worth reading are the ones that make the tier split a property of the system rather than a convention in ClojureScript — the server recomputes every tier-2 key it is handed, and refuses a block whose analysis does not declare a detector version.

And the browser one

Step 5's done-criterion is a PICTURE, and no assertion in cljs.test can check one: a take that resolves to the right numbers and draws nothing would pass every test in arthur.flow.freeze-test. A blank canvas under a perfectly correct transport is the bug class unit tests miss, and it has happened here once.

Step 9's is a picture too, for a different reason: the ways a document survives a round trip LOOKING correct are the interesting ones. So the suite now also saves the take, reopens it, and checks the frames are the same pixels.

It drives a real Chrome over CDP, and needs both processes up:

mise exec -- python manage.py runserver 8778              # from the REPO ROOT
cd frontend && mise exec -- npx shadow-cljs watch app     # in one shell
cd frontend && mise exec -- npm run browser               # in another

ARTHUR_URL overrides the page it drives; it defaults to http://localhost:8778/index.html, which since step 9 is Django's.

No dependencies. Playwright is not installed and CDP needs none — node --experimental-websocket has a global WebSocket and --headless=new --remote-debugging-port=N is the whole of the other side. It reads the canvas's own pixels rather than a screenshot, because the CSS scales the stage up by 2 and a screenshot is four pixels per raster pixel; it writes PNGs into test/browser/out/ anyway, so "it drew something" can be checked by eye as well as by count.

The app

Two processes, which do not talk to each other:

mise exec -- python manage.py runserver 8778              # from the REPO ROOT
cd frontend && mise exec -- npx shadow-cljs watch app

Then open http://localhost:8778/. Django serves the page from clips/templates/clips/index.html, and staticfiles serves the bundle out of static/arthur/js, where shadow-cljs already writes it — so nothing copies files between the two.

/index.html still works, and that is deliberate: it is the URL the browser suite has used since step 5, when shadow-cljs's :dev-http did no directory-index resolution and the suite learned to ask for the file.

Four built-in clips, on buttons in the transport:

take the synthetic take, head as filmed. Step 5's deliverable: a moving mouth, frozen into dense channels, with no video file anywhere.
locked the same freeze, head locked. The same blocks — :head's channels are written as framed identity instead of as a dense track, and nothing in tier 2 differs.
demo the hand-written scene from step 2. Not a face: the smallest scene that exercises every mechanism the model claims to have, so that each one is visible when it breaks.
swarm a hundred and twenty dense nodes. Not useful; it is the load test.

take and locked are the pair worth looking at together, because switching between them is the whole of what "stabilisation is a channel, not a mode" means.

The demo scene itself is src/arthur/demo/scene.edn. Both the synthetic take and real footage use src/arthur/flow/take.cljs for the measurement order and src/arthur/flow/freeze.cljs for the landmark-to-channel conversion.

Real footage

Choose a video in the footage file input. The server probes it, re-encodes it to a browser-seekable H.264 proxy, pulls WAV audio and one tracing JPEG per frame, then makes the resulting footage selectable. Click load frames to detect and freeze it. Extraction progress is currently read from /api/extractions/<key>; a future WebSocket can push the same job state. The uploaded bytes, extraction job, and decoded footage have separate records, so the same uploaded video can be reopened without decoding it again.

The proxy is what gets measured, and the stills are not. flow/ingest steps the proxy one frame at a time — seek to (i + 0.5) / fps, wait for requestVideoFrameCallback, check the mediaTime it reports is the frame that was asked for — and flow/detect hands each frame to MediaPipe in VIDEO running mode at i * 1000 / fps milliseconds. That timestamp has to increase strictly and has to be real footage time: video mode is a tracker, it reads the gap between timestamps as motion, and a repeat leaves the graph in an error state that every later call re-throws. The JPEGs beside the proxy are reference images for the tracing editor and nothing measures them, so they are not in the footage digest.

It is re-encoded even when the upload is already H.264, for two reasons: an iPhone's HEVC is not decodable in every browser, and the footage's identity is the proxy's digest — one produced by one ffmpeg invocation, not one that depends on which branch the source happened to take.

The command-line route is also available for an existing extracted bundle:

./extract.sh /path/to/clip.mov          # decode to frames + audio + manifest
mise exec -- python manage.py ingest_bundle

extract.sh keeps every source frame and writes frames/0001.png onward, audio.wav and manifest.json. Variable frame rate sources are rejected until the manifest and clock carry per-frame timestamps.

ingest_bundle then hashes all of it into the content-addressed blob store under var/blobs — by hard link, so 112MB of PNGs is not copied — and registers one Footage row. From then on the frames are the backend's: GET /api/footage/<id> answers with a manifest carrying a URL per frame, and the app fetches those.

That replaced a shared secret. Until step 9 the page fetched /manifest.json off the filesystem and built frames/0001.png itself, with shadow-cljs serving the repo root — so the frame layout was agreed between a shell script and a ClojureScript namespace, and "where are the frames" was answered by a directory listing. The cache-busting ?v= that used to hang off every frame URL went with it: a blob's name is the hash of its bytes, so re-extracting gives a frame a different URL rather than overwriting one.

To keep several takes, pass a bundle directory; each ingests separately and both stay selectable in the app.

./extract.sh /path/to/clip.mov scratch/my-take
mise exec -- python manage.py ingest_bundle scratch/my-take

scratch/ is ignored by Git, as are frames/, audio.wav and manifest.json at the root — all of it is extraction output, and tier 3 does not belong in the repo.

Loading detects one face per frame, measures the mouth, eyes and brows from landmarks and the teeth from source pixels, then freezes them into channels, and adds a button for the footage clip. Detection happens once when you load; playback only resolves channels and paints. Frames without a detection remain marked absent even though their neighbouring poses are used to condition the track. The scene now records stable subject and feature IDs and explicit eye pairs; dense channels can mark one feature absent while another is observed. Current MediaPipe loading supplies only the full-face detection mask. The stage stays 320×200 regardless of the footage dimensions. Real footage starts at the source picture rate. The picture fps buttons sample the frozen roto at lower rates while the source track, duration and audio clock stay unchanged. Picking frames to trace into cels is a separate future editing step. save also stores the detection mask, dense landmarks and raw RGBA mouth crops as three analysis blocks. open restores these without running MediaPipe or loading source PNGs. The frozen shapes remain separate channel blocks.

For known occlusion intervals, an extracted manifest may add "feature-absence": {"eye-r": [[10, 14]]}. Frame numbers are one-based and inclusive, matching PNG filenames. The eye remains the same feature when it reappears; the other eye and the mouth continue through the gap. This is an input annotation, with no UI for editing it yet.

MediaPipe's JS, wasm and model are under public/mediapipe/, served by Django's staticfiles under /static/mediapipe/. No CDN is used by this app. See that directory's README for provenance.

The server reports what it serves at GET /api/detector: the package version plus the sha256 of the model asset, and that string goes inside the content address of every block a detection produces. Asked rather than assumed, because a version constant in the client is one somebody has to remember to bump — and docs/architecture.md is explicit that a model upgrade silently reusing old landmarks presents as "the tool got worse", with no event to attach it to.

Port 8778 is deliberately not 8777. python3 serve.py from the repo root still runs the old JS tool on 8777, and the two are meant to run side by side.

Saving

save and open in the transport. A save has three ordered stages: is the tier split:

  1. the analysis record, so every block stored afterwards can name the detector version that produced it. The server refuses a block whose analysis it does not know.
  2. ask which blocks are missing, upload the source analysis blocks and frozen channel blocks, then link the source blocks to the analysis.
  3. the document — tier 1, as leaves. The server refuses a clip that names blocks it does not hold, so a saved document cannot load into a blank stage somewhere else.

The status line says what happened: saved r3 · 64 leaves · 8 blocks. Saving an unchanged document says 0 leaves · 0 blocks, which is both halves of the addressing working at once — an unchanged leaf keeps its version, and a content-addressed block is already there.

Two things are deliberately visible as failures. Saving swarm is refused, because its blocks have hand-written names and a document may only name content addresses. And open takes the most recently updated project and shows its first clip: there is no project browser, and the store holds one clip at a time.

The oracle, which is finished

js/ was the numeric oracle through step 4: test/parity/ ran both implementations on the same synthetic track and diffed fit-similarity, procrustes-mean, the raster and stabilize to 1e-9.

It was deleted at step 5, on purpose. Parity proves the port is FAITHFUL, not that the answer is RIGHT. The JS is a prototype and several of its conclusions contradict each other; a parity test pins behaviour while code moves, and keeping it afterwards would bake the prototype's mistakes into the rewrite and make them permanent. docs/port-plan.md says to delete it in one commit once the CLJS player renders the synthetic take, and that is what happened.

js/ itself stays as the reference for the MediaPipe setup, face measurements and pixel extraction. Its comments encode bugs that actually happened.

Layout

src/arthur/domain/      pure. No re-frame, no DOM, no flow/.
src/arthur/fx/          the only namespaces that talk to the network
src/arthur/flow/        the stages. `(f params inputs) -> output`, no state.
src/arthur/synth.cljs   the synthetic track. In src/ because the take PLAYS it —
                        it stands in for flow/detect, and a tool that needs a
                        video file before it shows you anything is one you
                        cannot debug.
src/arthur/demo.cljs    the hand-written scene, read from demo/scene.edn
src/arthur/demo/take.cljs  the synthetic source for the shared flow/take path
src/arthur/ui/canvas.cljs  indexed raster blit to the display canvas
test/arthur/support/    machinery shared between suites; not tests itself
test/browser/           drives a real Chrome over CDP. Not run by `npm test`.
public/mediapipe/       vendored wasm and model, served under /static/mediapipe/

public/ holds nothing but those assets now. The host page that used to sit beside them is clips/templates/clips/index.html.

Two evaluators, on purpose

domain/timeline has both eval-frame and resolver, and they are not alternatives:

  • (eval-frame timeline f store) is the specification. Allocating, order-free, obviously correct. Tests and one-off renders use it.
  • (resolver timeline store) -> (fn [f] ops) is what playback uses. It caches the topological order and the z paths, holds a cursor per channel and reuses one point buffer per node, so a frame allocates the op maps and nothing else.

Both run the same walk, parameterised by how a channel is read and where its points are written — two independent implementations of frame evaluation would drift, and the drift would look like a rendering bug rather than like two functions disagreeing. What differs between them is exactly the part that can be wrong, and scene-test asserts they agree frame for frame in forward, backward and random order.

Because the resolver reuses its buffers, ops must be rasterised before the next frame is asked for. That is the contract the rAF loop wants anyway: it reads, blits, and dispatches nothing.