Commit graph

16 commits

Author SHA1 Message Date
Olive Vaughn
179770d7d4 Split the shell into topbar, pool, stage, timeline and params panes
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 12:25:46 -04:00
Olive Vaughn
a45e89f4e4 Add polygon painting with per-gap drawing keys 2026-09-28 22:37:23 -04:00
Olive Vaughn
ffb95543a3 Add instanced 8625 stage with independent audio controls 2026-09-28 15:05:04 -04:00
Olive Vaughn
65ad67c129 Decode uploaded footage in order with WebCodecs 2026-09-28 14:18:09 -04:00
Olive Vaughn
83d106bbc5 Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.

Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:

/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.

A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.

Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.

Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.

The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.

Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
Olive Vaughn
686f897401 Add video upload, extraction progress, and reusable analysis sources 2026-09-28 10:45:51 -04:00
Olive Vaughn
9778b9023b Split clips from timelines 2026-09-28 02:33:26 -04:00
Olive Vaughn
9cd5243983 Serve the document from a Django backend, split into three tiers
Step 9. The tier split was the work; Django was the easy half.

Tier 1 — the authored scene — is the document, and it is addressed as
independently versioned leaves rather than saved whole, so one vertex drag
cannot clobber a collaborator's keying. `domain/leaf` is the document as
path -> value; `domain/wire` puts it on the wire as transit, because JSON
has neither integer map keys nor keywords and a save would quietly turn
`{0 v}` into `{"0" v}`.

Tier 2 — the dense channel blocks — is content-addressed by a hash over
every input, with the detector version inside every key through the
analysis the block descriptor names. `flow/address`'s `block-knobs` is the
invalidation table, and `address-test` does not trust it: it re-freezes the
take once per knob and asserts the biconditional, that a block's bytes
changed if and only if its key changed. That found `brow-pos` not depending
on `contour-avg` — the brow ring is smoothed, the raise is not.

Tier 3 — frames and audio — is served by the hash of its bytes out of the
same store. A manifest now names frames and carries a URL for each, so the
frame layout stopped being a shared secret between a shell script and a
ClojureScript namespace, and the `?v=` cache-buster went with it: a blob's
name is the hash of its contents, so a stale copy is not a thing that can
happen. The synthetic take's `audio.wav` moved to `static/arthur/` — an
asset the project owns, not an extraction that churns.

The server verifies rather than trusting a name it was handed: it
recomputes every key from the descriptor stored beside it, refuses an
analysis that declares no detector version, and refuses a document naming
blocks it does not hold. It hashes the descriptor TEXT, because JS prints
an integral double as `1` and Python as `1.0`, and a scheme where both ends
re-render the numbers disagrees on the first parameter that happens to be
whole.

Two loose ends from step 8 closed on the way. `pack` no longer takes a
`(track, frame)` predicate whose call sites each re-derived a feature from
an index — every track names the feature it follows, which deleted five
hand-maintained mappings. And `:dev-http` is gone: Django serves the page,
shadow-cljs only builds into the staticfiles tree.

227 CLJS tests, 31 Django tests, green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 01:11:41 -04:00
Olive Vaughn
35ef150b48 Give features stable identity, eye pairs and per-feature presence
Step 8's data model, ahead of its controls. Nothing here is a UI.

domain/params holds every knob's definition once — default, applicable area,
value constraints and the areas a change would force to regenerate. flow/take's
literal knob map becomes a view of it, so the take's defaults and the future
parameter panel cannot drift apart.

domain/feature adds subjects, features and groups as document data the renderer
never reads. A feature ID is stable for the whole clip, across occlusion: a run
of visible frames is not a new identity. An eye pair is an explicit group of one
or two eyes of the same subject, so a profile view with one identified eye needs
no invented partner. Settings resolve area -> subject -> group -> feature, and
dropping an eye from a pair materialises its effective values first so playback
does not jump. scene/problems now validates all of it.

Presence becomes per-feature rather than per-subject. freeze's :absent predicate
takes a track as well as a frame, so one occluded eye can be absent while its
partner still has a value; a full-face miss still marks everything absent. A
manifest may annotate known gaps as one-based inclusive intervals, which ingest
expands into observation tracks before measurement. An unobserved eye then gets
no vote in the iris pairing and cannot steer the shared gaze — gaze falls back to
whichever eye is visible. Temporal filters still see a sample on every frame,
held from the last observed one, because the numbers are a rectangular buffer;
the state mask, not the buffer, is what says the frame has no value.

js/app.js gets the same occlusion lesson: leading nulls from a face that starts
occluded used to throw away the whole take, and the neutral frame could be chosen
from a held duplicate pose.

Parameter editing, scoped regeneration and a feature-level detector remain. Until
one exists, footage without annotations falls back to the full-face mask rather
than claiming occlusions it cannot see.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B87NVmiU36qQmN9gmFYnJ9
2026-09-27 22:44:36 -04:00
Olive Vaughn
663b7c367a Add eyes brows and pixel-derived mouth interior to CLJS take 2026-09-27 19:48:42 -04:00
Olive Vaughn
06cf02db83 Preserve source cadence and sample picture fps after analysis 2026-09-27 19:31:39 -04:00
Olive Vaughn
32683efccf Port step 6: detect real footage with local MediaPipe 2026-09-27 19:13:20 -04:00
Olive Vaughn
8a06835895 Port step 5: freeze measured mouth into playable channels 2026-09-27 19:01:39 -04:00
Olive Vaughn
942e2f38ab Port step 4: measure the anchor and the mouth, condition on its own
`stabilize` is three things wearing one name, and it is now three functions in two
stages: `flow/measure/anchor` fits the rigid transform, `flow/condition` smooths
its parameters, `flow/measure/mouth` measures the lip rings through the result.
Parity is on the COMPOSITION and not on the pieces -- a split that agreed
function by function and not end to end would be a split rather than a port.

The oracle now drives `stabilize` at three configurations and the port agrees to
1e-9 on ref, rigid, transforms, outer, inner and aperture, plus `smoothContours`
at three radii. Two of the three configurations are at aspect 0.5625, a 1080x1920
phone clip, because at aspect 1 `pick` is the identity: a port that dropped the
anisotropy correction outright would pass every other assertion in the suite.
148 tests, up from 134.

Three decisions worth the reading time.

`makeXform` is not ported, and its absence takes the face oval with it. It
centres on the oval's bounding box and zooms until the face is 80% of the raster
height, so every vertex it touched carried a cropping decision made once, at
analysis time, from one frame's landmarks. Geometry belongs in the node's own
local space with the framing as a transform on a node, so this is a deletion. The
oval's only other consumer was the placeholder plate outline, which is painting.

The residual is taken against the RAW fit, and the prototype took it against the
smoothed one. That is the only deliberate numeric divergence here, and parity is
kept by asserting `anchor/residuals` on exactly what the prototype handed it. The
number's job is to say whether a section is stabilisable at all; folding the
smoothing error into it makes a slider look like a property of the footage, and
docs/architecture.md lists the residual under stage 3, which requires it to be
knob-free. `condition/anchor` therefore replaces `:transforms` and leaves
`:residual` alone.

The stage order is not the strict chain the table in docs/architecture.md looks
like, and that document now says so. The fit is knob-free, conditioning smooths
it, and the rings are measured *through* the conditioned transform -- so
`anchor avg` does re-run the ring mapping, which is a few hundred frames of twenty
points. The guarantee was only ever about the part that reads a source pixel, and
that part never sees a transform.

Two things fall out and are asserted rather than assumed. Smoothing and
subsampling commute, because both are per-slot, which is what lets `vertices`
stay a stage-5 knob downstream of a stage-4 one -- and it is also why the port can
smooth the full twenty slots where the prototype smooths eight and still match.
And `condition/contours` is `geom/moving-average` per vertex per axis rather than
its own clamped window, so "radius 2" cannot come to mean two different things at
the two knobs.

One dead end recorded so nobody walks it twice: the synth's head is perfectly
rigid -- its jitter is a whole-head translation, which a similarity absorbs
exactly -- so every frame's rigid configuration is congruent with frame zero's and
the Procrustes mean IS frame zero to 1e-15, jitter or none. "The reference is the
mean and not frame zero" cannot be asserted on this track and is asserted in
geom-test, where the two can differ.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-27 18:00:11 -04:00
Olive Vaughn
18d6495592 Port steps 2-3: the data model and the player
Steps 2 and 3 land together because the model revisions in the middle changed
code from both, and splitting them now would invent intermediate states that
never built.

  domain/channel  value-at across framed/keyed/dense, plus a cursor
  domain/node     decomposed transform, composition order, time maps
  domain/scene    topological order, z paths, eval-frame and resolver
  clock           audio-clocked frame derivation, outside app-db
  db/events/subs  re-frame arrives; the playhead is document state
  ui/player       the rAF loop; reads, blits, dispatches (almost) nothing
  ui/shell        transport

133 tests, 1158 assertions. The scene plays at 30fps against audio, scrubs, and
runs at 1/4x through 4x; verified by driving a real browser over CDP rather than
by assertion.

Two evaluators, on purpose. `eval-frame` is the specification -- allocating,
order-free, obviously correct. `resolver` is what playback uses: cached topo
order and z paths, a cursor per channel, a preallocated point buffer per node.
Both run the same walk, parameterised only by how a channel is read and where
points are written, because two independent implementations of frame evaluation
would drift and the drift would read as a rendering bug rather than as two
functions disagreeing. scene-test asserts they agree frame for frame in forward,
backward and random order.

Deviations and decisions, each with a reason:

- raster/fill-poly! is now a thin wrapper over fill-poly-buf!, which takes a flat
  preallocated buffer. ONE scanline fill serves the analysis stages, which speak
  {:x :y}, and frame evaluation, which hands over a buffer it owns. The parity
  suite still passes pixel-for-pixel, which is what makes the rewrite safe.

- The state mask carries ABSENCE ONLY. An earlier draft gave it a hidden bit too,
  per architecture.md's "hidden flag + palette index", and that bit was a dense
  [:vis] wearing a different hat -- two mechanisms for one question, which is how
  a part ends up hidden by one and shown by the other.

- The palette is a parameter of evaluation, not a global. A node names a TONE;
  which ramp that tone is read in belongs to the timeline it sits in.

- :over layers and a symbol :rate THROW rather than being ignored. Neither is
  built and nothing can produce one, so this can only fire on data that has run
  ahead of the code. A silently dropped override is a hand correction the user
  made once, watched fail, and has no reason to trust again.

Three findings the model produced rather than received:

- Presence propagates asymmetrically. An absent transform drops the subtree; an
  absent [:geom :pts] drops only that node, because an absent mouth outline has
  nothing to draw but the head it hangs off has not moved. That asymmetry is the
  reason presence is tracked per channel and not per node.

- Z paths need lexicographic compare, not `compare`, which orders vectors by
  count first -- so a cel three levels under "a1" would jump in front of a bare
  "a2" and the layer order would mostly work.

- A node stencilled by something that drew nothing is dropped, not drawn
  unclipped: an iris floating over the cheek is worse than a missing iris.

docs/ revised alongside, and those revisions are the load-bearing part:

- A scene, a timeline and a symbol are one type. The doc had two structures with
  the same fields and never said so. Two axes of nesting are now separated --
  parent/child within a timeline is flat with parent pointers, instance nesting
  is by reference -- which is why "nestable" and "flat" only sounded
  contradictory.

- Palettes are named, live on the project, and are ENABLED on a timeline as a
  channel. Absent inherits; present travels with the timeline, so a symbol
  authored against :night stays night wherever it is placed. The output index
  space is the concatenation of the named ramps, which keeps one buffer and one
  flat table and incidentally stops two nodes in different palettes colliding on
  a stencil.

- Stabilisation is a channel, not a mode: {s, theta, tx, ty} IS [:xform :*], so
  the normalise on/off/per-plate toggle is which of the three channel shapes the
  :head node carries. Always measure and always store factored -- smoothing and
  velocity-minimum key selection both need the split to exist in storage.

- There is no camera node and none is needed. Placement is a node transform, the
  stage clips what hangs off it, and project dimensions are independent of the
  footage. `makeXform` is therefore not to be ported: it bakes a cropping
  decision into every stored vertex.

- Export is removed. The .take writer was for an Animator Pro render script; the
  target is encoding video in the browser, and step 9 now says not to port the
  old one.

demo/swarm is 120 shapes on six orbits, entirely dense blocks behind store
handles -- the shape freeze produces at step 5, and the first thing to exercise
that path under load. It plays at 30fps, and bench-test keeps a deliberately
loose floor under it because a performance regression here does not announce
itself: the picture stays correct and merely arrives late.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PDfHGdV39zu6rvgbBTfDaT
2026-09-27 17:28:05 -04:00
Olive Vaughn
eb06be005c Port steps 0-1: scaffold, the oracle, and the pure bottom
Scaffolds frontend/ (shadow-cljs, reagent 1.2.0, re-frame 1.4.3) and ports
everything below the data model, with the JS kept as a numeric oracle.

  domain/landmarks  index tables, verbatim
  domain/ring       subsample, offset, simplicity
  domain/geom       similarity fit, procrustes, moving average
  domain/raster     indexed scanline fill, stencil, disc, rect
  domain/palette    the ramp, and the no-sampled-RGB rule

58 tests, 166 assertions. Parity with js/ on the identical 72-frame synthetic
track: fit-similarity, procrustes-mean, fit-residual, moving-average,
smooth-transforms, offset-ring and subsample-slots to 1e-9; the raster
pixel-for-pixel over the whole buffer.

Three deviations from the JS, each for a reason:

- synth.cljs jitters from a SEEDED generator, not Math.random. Parity is only
  checkable if both sides can be handed the same track, and a failing assertion
  has to be reproducible. `:rand-fn` takes the generator over, so oracle.mjs
  stubs js/Math.random and js/ itself stays untouched.

- raster/->rgba replaces toImageData. ImageData is a DOM type and domain/ may
  not touch the DOM; returning plain bytes also lets the
  no-intermediate-colours assertion run in node. ui/canvas wraps it later.

- offset-ring lives in domain/ring, not domain/geom, per architecture.md: it is
  an operation on an ordered traversal, not on a transform.

Step 1's "done" also names the swapped-iris vote, but pairIrises is in
pipeline.js and belongs to step 7. The precondition is asserted instead --
`:swap-iris` really does move both blocks -- so the vote will have a track that
disagrees with it when it arrives.

One finding, recorded in full in the test that measures it: smooth-transforms
buys nothing on the synthetic track. Against jitter-free ground truth, radius 1
helps by 17% on one noise realisation and hurts by 0.5% on another, so its
benefit is within noise; from radius 2 up the cost is unambiguous, and by radius
5 the filter is below the true motion's own high-frequency energy, i.e.
smoothing away performance. The test pins the shape of the knob rather than a
preferred value. This may say more about the synth's jitter being unrealistically
small (+/-0.001 normalised) than about the knob; step 6 settles it on real
footage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-27 14:43:34 -04:00