Preserve source cadence and sample picture fps after analysis

This commit is contained in:
Olive Vaughn 2026-09-27 19:31:39 -04:00
parent 32683efccf
commit 06cf02db83
20 changed files with 303 additions and 87 deletions

2
.gitignore vendored
View file

@ -1,4 +1,6 @@
frames/
# local extracted takes for comparing source cadences
/scratch/
*.task
*.take

View file

@ -42,7 +42,7 @@ wasm, which is fetched from a CDN on first use.
For real footage:
```sh
./extract.sh /path/to/clip.mov 12 # -> frames/*.png, audio.wav, manifest.json
./extract.sh /path/to/clip.mov # -> frames/*.png, audio.wav, manifest.json
```
then **Load frames**. MediaPipe's wasm is fetched from jsdelivr on first use;
@ -52,16 +52,16 @@ Frames are pre-extracted rather than decoded in the page because browser video
seeking is approximate and `requestVideoFrameCallback` only delivers frames at
playback speed — neither gives a deterministic per-frame pass.
`manifest.json` records the true extraction rate. The page reads it rather than
assuming, because a guessed fps desynchronises audio from picture — and sync is
the one thing this view exists to show.
`manifest.json` records the source rate. The extractor keeps every source frame;
the page reads that rate because a guessed fps desynchronises audio from picture.
Choosing a lower picture rate happens after analysis.
**Exposure** decides how often the picture gets a new drawing: rip at 24 and
render `on 2s` for 12, `on 3s` for 8. The dense track and the audio are
untouched, so it is a dropdown rather than a re-rip, and the export emits keys
only on the grid instead of the same pose twice. Everything rides the same grid
— mouth, eyes, teeth, plate — because a head cutting on the odd frames while the
mouth cuts on the even ones reads as two performances laid over each other.
**Picture fps** decides how often the finished roto gets a new pose. Analyze all
source frames, then sample those frozen poses at 12, 24 or the source rate while
keeping the original duration and audio. **Exposure** can hold a drawing across
more than one picture slot. The tracing editor chooses source frames for cel
references separately. Shared timing is the useful default for mouth, eyes,
teeth and plate so their changes read as one performance.
**Audio is the playback clock**: `frame = floor(audio.currentTime * fps)`. A slow
render loop therefore drops frames instead of drifting, and ½x / ¼x work by
@ -322,13 +322,13 @@ the tool a person made by hand; everything else regenerates. They are not in the
## Two kinds of sparseness
Sparseness has two unrelated causes, and conflating them was the original design
error here. **Aesthetic** sparseness is set by the extraction rate — pick 12fps and
you have already chosen your timing. **Labour** sparseness is a human drawing
each one, and it binds only on the plate.
error here. **Aesthetic** sparseness is chosen from the full analyzed source
track at rendering time. **Labour** sparseness is a human drawing each cel and
selecting which source frames to use as tracing references.
Aesthetic sparseness is the **exposure** control, not the extraction rate —
making it a render-time grid means auditioning 12 against 24 costs a dropdown
instead of a re-rip and a full re-detection.
Aesthetic sparseness is the **picture fps** control, with exposure available for
longer holds. Both happen after analysis, so auditioning 12 against 24 needs no
re-extraction or re-detection.
So the mouth keeps **every** frame: it is traced, and therefore free. In limited
animation lip sync is routinely the densest element, on 1s, while heads hold on

View file

@ -333,11 +333,15 @@ Every node may map the frame it is evaluated at:
```clojure
:time {:mode :inherit} ; the default, and almost always right
:time {:mode :map :expose 2 :offset -1 :rate 1.0 :loop? false}
:time {:mode :map :source-fps 30 :sample-fps 12} ; root: lower picture cadence
```
Three features that look unrelated are this one mechanism:
- **exposure** is `⌊f/n⌋·n`,
- **picture fps** quantises source time to a chosen picture grid, then reads the
latest source pose at or before that time; source analysis and audio keep their
original cadence,
- **mouth lead** is `f + k`,
- **a symbol instance's timing** is `(f - at)·rate + in`, with optional looping.
@ -463,8 +467,11 @@ world transform of the node it rides:
Composed with image-pixels-to-local — **both axes divided by `imgH`**, never by
their own dimension — the photo is registered with the shapes by construction,
and an unregistered underlay is merely decorative. Which frame it shows, the
current one or the held plate frame, is a UI choice and not a stored one.
and an unregistered underlay is merely decorative. The tracing editor chooses
which source frame to show under a cel. That reference choice is independent of
the finished picture fps and does not change the dense analysis track. A cel can
therefore use any useful source frame as its drawing reference, even when that
frame is not one of the displayed picture poses.
A photo that has to sit *between* two drawn layers is the case that would make it
a `:bitmap` node with an op of its own. Nothing wants that yet: a reference is
@ -527,6 +534,7 @@ Proof that it covers what exists, not just what is wanted:
| painted background cel | node per layer, `[:geom :pts]` **framed**, `[:style :color]` framed |
| `mouth lead` | `:time {:offset k}` on performance nodes only |
| `exposure` | `:time {:expose n}` on the clip root, inherited |
| picture fps | `:time {:source-fps s :sample-fps p}` on the clip root, applied after analysis |
| hand correction | an `:over` layer, `:offset` or `:replace` |
The brow row is the one worth looking at twice. `docs/design.md` argues at length

View file

@ -49,6 +49,11 @@ overrides and kept-frame sets are all in clip-frame space, so a clip slides on
the timeline without a single stored number changing. Exposure and lead are
transforms *within* clip space:
Detection retains every source frame. A chosen picture fps samples the frozen
roto in clip time; it changes neither source-frame count nor the audio clock.
The set of source frames an artist uses as cel tracing references is another
selection, independent of the picture fps.
```clojure
(defn pose-frame [clip cf]
(-> cf (expose (:exposure clip)) (shift (:lead clip) (count-frames clip))))
@ -73,7 +78,7 @@ project
│ ├── scene node tree
│ ├── channels node+property -> keyframe stream, in cf
│ ├── overrides (node, property, cf) -> value, applied last
│ └── cels painted vector layers, per exposure slot
│ └── cels painted vector layers, keyed by chosen clip frames
└── sequence[] clip placements: {clip-id, at, in, out}
```
@ -81,6 +86,10 @@ project
it a name is most of the work: `state` in `app.js` is a clip with its analysis
inlined and its palette global.
Cel keys select where drawings begin and how long they hold. The source frames
shown beneath a cel while tracing are chosen independently, and picture fps
only controls which analyzed pose the finished roto displays at a given time.
### Two things called "track"
`docs/design.md` says "dense track" for the landmark stream. A timeline also

View file

@ -4,8 +4,10 @@ Self-contained. You should not need any prior conversation to execute this.
**Implementation status (2026-09-27):** steps 0–6 are in the CLJS frontend.
Step 6 reads extracted footage from the manifest, detects landmarks with local
MediaPipe assets, and runs the same freeze path as the synthetic take. Step 7 is
next: eyes, brows and the pixel-derived mouth interior.
MediaPipe assets at full source cadence, and runs the same freeze path as the
synthetic take. The scene time map can sample the frozen roto at a lower picture
fps without changing source analysis, duration or audio. Step 7 is next: eyes,
brows and the pixel-derived mouth interior.
## What arthur is
@ -302,6 +304,10 @@ MediaPipe interop behind one namespace; real frames, real audio, real fps from
the manifest. **Vendor the wasm** rather than fetching from jsdelivr — it is
currently the only thing in the tool that silently requires a network.
Decode every source frame for analysis. A lower output picture fps is a time map
over frozen channels, not a reduced detection track. Selecting source frames to
trace into cels is independent again and remains outside this port's paint scope.
**Done:** real footage plays back as a rotoscoped mouth.
### 7 — the rest of measure

View file

@ -7,32 +7,77 @@
# pass. A PNG sequence is exact, instantly seekable, and reproducible.
#
# Audio comes out alongside because the page uses it as the PLAYBACK CLOCK -
# frame = floor(audio.currentTime * fps) - so picture and sound cannot drift
# apart no matter how long the shot is or how slow the render loop runs.
# frame = floor(audio.currentTime * fps). Detection sees every decoded source
# frame; a lower drawing rate is a later playback choice, never an extraction
# choice. This script accepts CFR footage because frame-index timing needs a
# single rate. VFR needs per-frame timestamps in the manifest first.
set -euo pipefail
src="${1:?usage: ./extract.sh CLIP [FPS] [OUTDIR]}"
fps="${2:-12}"
out="${3:-frames}"
rm -rf "$out"; mkdir -p "$out"
ffmpeg -hide_banner -loglevel warning -i "$src" -vf "fps=$fps" "$out/%04d.png"
count=$(ls -1 "$out" | wc -l)
# Mono is enough for judging sync and halves the file. Absent audio is not fatal.
if ffprobe -v error -select_streams a:0 -show_entries stream=codec_type \
-of csv=p=0 "$src" 2>/dev/null | grep -q audio; then
ffmpeg -hide_banner -loglevel warning -y -i "$src" -vn -ac 1 -ar 44100 audio.wav
audio='"audio.wav"'
echo "audio -> audio.wav"
else
audio='null'
echo "no audio stream"
src="${1:?usage: ./extract.sh CLIP [BUNDLE_DIR]}"
bundle="${2:-.}"
if [[ "$bundle" =~ ^[0-9]+([.][0-9]+)?$ ]]; then
echo "The FPS argument was removed: extraction always keeps the source rate. Use a directory as argument 2." >&2
exit 2
fi
if [[ "$bundle" = /* || "$bundle" = *..* ]]; then
echo "BUNDLE_DIR must be a relative directory inside this repo" >&2
exit 2
fi
# The page must know the true extraction rate: if it guessed, audio and picture
# would drift. Source of truth lives here, next to the frames it describes.
printf '{"fps":%s,"frames":%s,"dir":"%s","audio":%s,"source":"%s"}\n' \
"$fps" "$count" "$out" "$audio" "$(basename "$src")" > manifest.json
probe=$(ffprobe -v error -select_streams v:0 \
-show_entries stream=r_frame_rate,avg_frame_rate,nb_frames \
-of json "$src")
fps=$(python3 -c '
import json, sys
from fractions import Fraction
streams = json.load(sys.stdin).get("streams", [])
if not streams:
raise SystemExit("no video stream in source")
s = streams[0]
nominal = Fraction(s["r_frame_rate"])
average = Fraction(s["avg_frame_rate"])
if nominal <= 0 or average <= 0 or abs(float(nominal / average) - 1) > 0.001:
raise SystemExit("variable-frame-rate source needs timestamp-aware playback; refusing to guess its fps")
print(float(average))
' <<< "$probe")
echo "$count frames at ${fps}fps -> $out/ (manifest.json written)"
if [[ "$bundle" = "." ]]; then
dir="frames"; audio_path="audio.wav"; manifest_path="manifest.json"
else
dir="${bundle%/}/frames"
audio_path="${bundle%/}/audio.wav"
manifest_path="${bundle%/}/manifest.json"
fi
rm -rf "$dir"; mkdir -p "$dir"
ffmpeg -hide_banner -loglevel warning -i "$src" -fps_mode passthrough "$dir/%04d.png"
count=$(find "$dir" -maxdepth 1 -name '*.png' -type f | wc -l | tr -d ' ')
expected=$(python3 -c 'import json,sys; print(json.load(sys.stdin)["streams"][0].get("nb_frames", ""))' <<< "$probe")
if [[ "$expected" =~ ^[0-9]+$ && "$count" != "$expected" ]]; then
echo "decoded $count frames but source reports $expected; refusing an inaccurate manifest" >&2
exit 1
fi
# Mono is enough for judging sync and halves the file. A silent clock lets a
# mute source use the same audio-driven transport.
if ffprobe -v error -select_streams a:0 -show_entries stream=codec_type \
-of csv=p=0 "$src" 2>/dev/null | grep -q audio; then
ffmpeg -hide_banner -loglevel warning -y -i "$src" -vn -ac 1 -ar 44100 "$audio_path"
else
duration=$(python3 -c 'import sys; print(int(sys.argv[1]) / float(sys.argv[2]))' "$count" "$fps")
ffmpeg -hide_banner -loglevel warning -y -f lavfi -i anullsrc=r=44100:cl=mono \
-t "$duration" -c:a pcm_s16le "$audio_path"
fi
# JSON escaping belongs to a JSON writer, especially for source filenames.
python3 - "$fps" "$count" "$dir" "$audio_path" "$src" "$manifest_path" <<'PY'
import json, os, sys
fps, count, frames, audio, source, path = sys.argv[1:]
with open(path, "w") as out:
json.dump({"fps": float(fps), "frames": int(count), "dir": frames,
"audio": audio, "source": os.path.basename(source)}, out)
out.write("\n")
PY
echo "$count source frames at ${fps}fps -> $dir/ ($manifest_path written)"

View file

@ -87,16 +87,32 @@ and real footage use `src/arthur/flow/take.cljs` for the measurement order and
From the repo root, extract a clip, then click **load frames** in the CLJS app:
```sh
./extract.sh /path/to/clip.mov 12
./extract.sh /path/to/clip.mov
```
This writes `frames/0001.png` onward, `audio.wav`, and `manifest.json` at the
repo root. The manifest supplies the exact frame count, fps and audio path.
This keeps every source frame and writes `frames/0001.png` onward, `audio.wav`,
and `manifest.json` at the repo root. The manifest supplies the exact frame
count, source fps and audio path. Variable frame rate sources are rejected until
the manifest and clock carry per-frame timestamps.
To keep multiple takes or compare with a previous extraction, pass a bundle
directory and enter its manifest path in the app:
```sh
./extract.sh /path/to/clip.mov scratch/my-take
# source manifest field: /scratch/my-take/manifest.json
```
`scratch/` is ignored by Git. The directory contains its own frames, audio and
manifest, so extracting it does not replace another take's files.
Loading detects one face per frame, freezes the measured mouth into channels,
and adds a button for the footage clip. Detection happens once when you load;
playback only resolves channels and paints. Frames without a detection remain
marked absent even though their neighbouring poses are used to condition the
track. The stage stays 320×200 regardless of the footage dimensions.
track. The stage stays 320×200 regardless of the footage dimensions. Real
footage starts at the source picture rate. The **picture fps** buttons sample the
frozen roto at lower rates while the source track, duration and audio clock stay
unchanged. Picking frames to trace into cels is a separate future editing step.
MediaPipe's JS, wasm and model are under `public/mediapipe/` and served locally.
No CDN is used by this app. See that directory's README for provenance.

View file

@ -31,6 +31,12 @@
.scrub { width: 100%; margin: 10px 0 6px; }
.readout { display: flex; gap: 18px; opacity: .55; font-size: 12px; }
.readout .warn { color: #d98f5a; opacity: 1; }
.picture-rate { display: flex; align-items: center; gap: 6px; margin-top: 7px;
font-size: 12px; }
.source-path { display: block; margin-top: 8px; font-size: 12px; opacity: .7; }
.source-path input { width: 300px; margin-left: 8px; padding: 3px 5px;
color: var(--fg); background: #1c1f2b; border: 1px solid #2b3040;
font: inherit; }
.load-status { margin-top: 6px; font-size: 12px; opacity: .75; }
.note { opacity: .35; font-size: 12px; max-width: 640px; }
</style>

View file

@ -97,3 +97,8 @@
rather than only inside the scene."
[f expose]
(node/expose f expose))
(defn picture-frame
"The source pose displayed at f after picture-rate sampling and exposure."
[f source-fps picture-fps expose]
(node/expose (node/sample-frame f source-fps picture-fps) expose))

View file

@ -23,7 +23,8 @@
because there is one clip per scene today. Copying them by hand into this table
is how one of them comes to disagree with the scene it describes."
[label scene store]
(merge {:label label :scene scene :store store :audio "/audio.wav"}
(merge {:label label :scene scene :store store :audio "/audio.wav"
:display-fps (:fps scene)}
(select-keys scene [:fps :frames :width :height])))
(def scenes
@ -50,9 +51,10 @@
;; footage's. That is what deleting `makeXform` buys — the framing became a
;; transform on a node, so nothing downstream of the freeze knows the frame
;; size — and it is why ui/player no longer hardcodes 320x200.
:clip (select-keys (:take scenes) [:fps :frames :width :height :audio])
:clip (select-keys (:take scenes) [:fps :frames :width :height :audio :display-fps])
:footage {:id nil :label nil :loading? false :status nil}
:footage {:id nil :label nil :loading? false :status nil
:manifest-path "/manifest.json"}
;; --- transport ---
;;

View file

@ -85,6 +85,22 @@
[f n]
(if (and n (> n 1)) (* (js/Math.floor (/ f n)) n) f))
(defn sample-frame
"Pick a source frame for a lower picture rate without changing clip time.
The input is already an integer source frame from the audio clock. Its time is
f/source-fps. Quantise that time to the picture grid, then read the latest
source frame at or before it. The result is always an integer and never from
the future, including when the rates do not divide (30 source → 24 picture)."
[f source-fps picture-fps]
(if (and source-fps picture-fps
(pos? source-fps) (pos? picture-fps)
(< picture-fps source-fps))
(min f (js/Math.floor
(* (js/Math.floor (/ (* f picture-fps) source-fps))
(/ source-fps picture-fps))))
f))
(defn local-frame
"Apply a node's time map to the frame it was handed by its parent.
@ -99,7 +115,8 @@
as two performances; offset is PER-NODE by design, because mouth lead applies
to performance nodes and not to the plate, which is the entire point of it."
[n f]
(let [{:keys [mode offset rate] ex :expose :or {mode :inherit}} (:time n)]
(let [{:keys [mode offset rate source-fps sample-fps]
ex :expose :or {mode :inherit}} (:time n)]
(if (= mode :inherit)
f
(do
@ -110,7 +127,11 @@
(when (and rate (not= rate 1.0) (not= rate 1))
(throw (ex-info "time map :rate is symbol timing and symbols are not built (port-plan step 2 scope)"
{:node (:id n) :time (:time n)})))
(when (and sample-fps (not (and source-fps (pos? source-fps))))
(throw (ex-info "picture sampling needs a positive source fps"
{:node (:id n) :time (:time n)})))
(cond-> f
sample-fps (sample-frame source-fps sample-fps)
ex (expose ex)
offset (+ offset))))))
@ -256,4 +277,3 @@
(into (for [[path c] (:channels n)
p (ch/problems c)]
(str "channel " (pr-str path) ": " p))))))

View file

@ -43,13 +43,11 @@
(defn- build-clip [manifest {:keys [dense detected dimensions missing first-real]}]
(let [[w h] dimensions
params (merge take/knobs
{:name "footage" :fps (:fps manifest) :aspect (/ w h)
:stage [320 200] :expose 2 :head :as-filmed
:analysis (str "mediapipe:1.0.1/" (:source manifest))})
frozen (take/build params {:dense dense :detected detected})
frozen (take/footage manifest {:dense dense :detected detected
:dimensions dimensions})
scene (:scene frozen)]
(assoc (select-keys scene [:fps :frames :width :height])
:display-fps (:fps scene)
:scene scene :store (:store frozen)
;; Re-extraction often overwrites audio.wav under the same name. A new
;; URL makes the element fetch the new sound when this clip is loaded.
@ -64,8 +62,8 @@
(rf/reg-fx
::begin!
(fn [_]
(-> (ingest/manifest!)
(fn [path]
(-> (ingest/manifest! path)
(.then (fn [manifest]
(rf/dispatch [::progress "loading MediaPipe…"])
(-> (detect/landmarker!)
@ -88,7 +86,11 @@
{:db (assoc db :footage (assoc (:footage db) :loading? true
:status "reading manifest.json…"))
::pb/pause! nil
::begin! nil})))
::begin! (get-in db [:footage :manifest-path])})))
(rf/reg-event-db
::set-manifest-path
(fn [db [_ path]] (assoc-in db [:footage :manifest-path] path)))
(rf/reg-event-db
::progress
@ -106,9 +108,9 @@
(let [clip (store/entry id)]
{:db (-> db
(assoc :scene/current id
:clip (select-keys clip [:fps :frames :width :height :audio])
:footage {:id id :label (:label clip)
:loading? false :status summary})
:clip (select-keys clip [:fps :frames :width :height :audio :display-fps])
:footage (assoc (:footage db) :id id :label (:label clip)
:loading? false :status summary))
(assoc-in [:playback :frame] 0)
(assoc-in [:playback :playing?] false))
::pb/pause! nil})))

View file

@ -65,6 +65,13 @@
{:db (assoc-in db [:playback :rate] r)
::rate! r}))
(rf/reg-event-db
::set-picture-fps
(fn [db [_ target]]
(if (and (number? target) (pos? target) (<= target (fps db)))
(assoc-in db [:clip :display-fps] target)
db)))
;; --- effects: every DOM touch on the audio element is one of these ---
(rf/reg-fx ::play! (fn [_] (clock/play!)))
@ -98,7 +105,7 @@
;; The stage travels with the clip: two clips may be different
;; sizes, and the raster the loop paints into is the clip's, not
;; the app's.
(assoc :clip (select-keys clip [:fps :frames :width :height :audio]))
(assoc :clip (select-keys clip [:fps :frames :width :height :audio :display-fps]))
(assoc-in [:playback :frame] 0)
(assoc-in [:playback :playing?] false))
::pause! nil

View file

@ -275,6 +275,31 @@
;; ---------------------------------------------------------------------------
;; the face, onto the stage
(defn- motion-placement
"An editable default framing for a real take. Fit the observed eye/nose and
mouth motion within the stage; otherwise a downward head move can push the
mouth below a fixed 320×200 canvas even while it stays in the source image."
[[w h] {:keys [rigid transforms outer detected]}]
(let [points (mapcat (fn [i]
(when (or (nil? detected) (nth detected i))
(concat (nth rigid i)
(geom/apply-sim-all (invert (nth transforms i))
(nth outer i)))))
(range (count outer)))
xs (map :x points)
ys (map :y points)
x0 (reduce min xs)
x1 (reduce max xs)
y0 (reduce min ys)
y1 (reduce max ys)
cx (/ (+ x0 x1) 2)
cy (/ (+ y0 y1) 2)
k (min (/ (* 0.8 w) (max 1e-9 (- x1 x0)))
(/ (* 0.8 h) (max 1e-9 (- y1 y0))))]
{[:xform :anchor] (ch/framed [cx cy])
[:xform :scale] (ch/framed [k k])
[:xform :pos] (ch/framed [(- (/ w 2) cx) (- (/ h 2) cy)])}))
(defn face-placement
"The face's transform on the stage, as FRAMED channels.
@ -284,9 +309,14 @@
ordinary node. `makeXform` made the same decision and then baked it into every
vertex, where nothing could ever revise it.
Derived from the reference rigid configuration, which is what the freeze has:
the face oval is not measured, because its only consumers in the prototype were
that transform and the placeholder plate outline. Two numbers:
The synthetic take uses the reference rigid configuration for its default.
Real footage can request `:fit-motion?`: its default fits the observed rigid
and lip motion in the stage so a head move does not send the mouth off canvas.
Both are ordinary editable transforms on :face, never baked into the geometry.
The face oval is not measured, because its only consumers in the prototype were
the old baked framing transform and the placeholder plate outline.
For the synthetic default, two numbers:
SCALE is stage pixels per image height, set so the reference's eye-corner span
is 40% of the stage width. Landmark-free — it is the rigid configuration's own
@ -304,15 +334,17 @@
landmarks are eyes and nose — the upper middle of a face — so a quarter down
leaves the jaw and the mouth on the stage. Whatever hangs off is clipped, which
is not a feature to add: every fill in `domain/raster` clamps already."
[{:keys [stage]} {:keys [ref]}]
(let [[w h] stage
c (geom/centroid ref)
span (- (reduce max (map :x ref)) (reduce min (map :x ref)))
k (/ (* 0.4 w) span)]
{[:xform :anchor] (ch/framed [(:x c) (:y c)])
[:xform :scale] (ch/framed [k k])
[:xform :pos] (ch/framed [(- (/ w 2) (:x c))
(- (* 0.25 h) (:y c))])}))
[{:keys [stage fit-motion?]} {:keys [ref] :as inputs}]
(if fit-motion?
(motion-placement stage inputs)
(let [[w h] stage
c (geom/centroid ref)
span (- (reduce max (map :x ref)) (reduce min (map :x ref)))
k (/ (* 0.4 w) span)]
{[:xform :anchor] (ch/framed [(:x c) (:y c)])
[:xform :scale] (ch/framed [k k])
[:xform :pos] (ch/framed [(- (/ w 2) (:x c))
(- (* 0.25 h) (:y c))])})))
;; ---------------------------------------------------------------------------
;; the clip
@ -327,6 +359,8 @@
`(count outer)`, because a freeze that could disagree with its
own input about the length of the take would.
:stage [w h] project dimensions, INDEPENDENT of the footage
:fit-motion? choose an editable real-footage default that keeps the
observed feature motion on stage
:expose the clip root's exposure grid, inherited by everything
:verts the lip rings' vertex budget
:aperture-cut fraction of the take's peak aperture below which the mouth

View file

@ -13,8 +13,8 @@
(assoc m :fps fps :frames frames)))
(defn manifest!
[]
(-> (js/fetch "/manifest.json" #js {:cache "no-store"})
[path]
(-> (js/fetch (str "/" (str/replace path #"^/+" "")) #js {:cache "no-store"})
(.then (fn [response]
(when-not (.-ok response)
(throw (ex-info "manifest.json was not found; run extract.sh first"

View file

@ -24,3 +24,16 @@
(defn build
[params inputs]
(freeze/clip params (measure params inputs)))
(defn footage
"A real manifest and its detected landmarks through the same measurement and
freeze path as the synthetic take. The source cadence stays in :fps; picture
sampling is a root time map applied only after this artifact exists."
[manifest {:keys [dense detected dimensions]}]
(let [[w h] dimensions
params (merge knobs
{:name "footage" :fps (:fps manifest) :aspect (/ w h)
:stage [320 200] :fit-motion? true
:expose 1 :head :as-filmed
:analysis (str "mediapipe:1.0.1/" (:source manifest))})]
(build params {:dense dense :detected detected})))

View file

@ -10,6 +10,7 @@
(rf/reg-sub ::loop? (fn [db _] (get-in db [:playback :loop?])))
(rf/reg-sub ::muted? (fn [db _] (get-in db [:playback :muted?])))
(rf/reg-sub ::fps (fn [db _] (get-in db [:clip :fps])))
(rf/reg-sub ::display-fps (fn [db _] (get-in db [:clip :display-fps])))
(rf/reg-sub ::frames (fn [db _] (get-in db [:clip :frames])))
;; The stage, in pixels. On the clip because project dimensions are independent
;; of the footage — see flow/freeze/face-placement — so the canvas and the raster

View file

@ -10,11 +10,28 @@
(:require [arthur.domain.palette :as pal]
[arthur.domain.scene :as scene]
[arthur.footage.store :as footage]
[arthur.subs.playback :as playback]
[re-frame.core :as rf]))
(rf/reg-sub ::scene-id (fn [db _] (:scene/current db)))
(rf/reg-sub ::scene (fn [db _] (:scene (footage/entry (:scene/current db)))))
(rf/reg-sub
::base-scene
:<- [::scene-id]
(fn [id _] (:scene (footage/entry id))))
(rf/reg-sub
::scene
:<- [::base-scene]
:<- [::playback/display-fps]
(fn [[scene picture-fps] _]
;; This changes only the scene's root time map. The dense source track stays
;; at its native rate, and the audio clock still advances through source time.
(if (and scene picture-fps (< picture-fps (:fps scene)))
(-> scene
(assoc-in [:nodes :root :time :source-fps] (:fps scene))
(assoc-in [:nodes :root :time :sample-fps] picture-fps))
scene)))
(rf/reg-sub
::exposure
@ -45,11 +62,12 @@
(rf/reg-sub
::store
(fn [db _]
:<- [::scene-id]
(fn [id _]
;; Tier 2, behind a handle, and never in app-db itself — what is in the db is
;; the id of the clip whose blocks these are. The hand-written demo has none;
;; the swarm is entirely dense.
(:store (footage/entry (:scene/current db)))))
(:store (footage/entry id))))
(rf/reg-sub
::resolver

View file

@ -167,12 +167,18 @@
;; the next one lands where the audio already is, so the failure mode is
;; a visible stutter rather than an invisible slide out of sync.
f (if live? (clock/frame fps frames) frame)]
;; A paused scrub is an explicit seek, not a playback sample. Clear the
;; meter before the next play so its first frame cannot count the seek as
;; dropped frames or leave a stale low rate in the transport.
(when (and (not live?) (:t0 @meter-state))
(reset! meter-state {:t0 nil :paints 0 :samples 0 :advanced 0 :prev nil})
(reset! meter {:fps 0 :drop 0}))
;; `ready?` gates the bookkeeping as well as the draw: recording a frame as
;; painted when it was not is how the canvas stays empty forever.
(when (and ready? f (not= f (:last @state)))
(swap! state assoc :last f)
(paint! f)
(meter! f)
(when live? (meter! f))
(when live? (rf/dispatch [::pb/tick f])))
;; The audio ending is the authority on playback having stopped; nothing
;; counts frames to notice it.

View file

@ -35,8 +35,10 @@
frame @(rf/subscribe [::sub/frame])
frames @(rf/subscribe [::sub/frames])
fps @(rf/subscribe [::sub/fps])
picture-fps @(rf/subscribe [::sub/display-fps])
current @(rf/subscribe [::render/scene-id])
expose @(rf/subscribe [::render/exposure])
{:keys [id label loading? status]} @(rf/subscribe [::sub/footage])]
{:keys [id label loading? status manifest-path]} @(rf/subscribe [::sub/footage])]
[:div.transport
[:div.row
[:button {:on-click #(rf/dispatch [::pb/toggle])}
@ -77,8 +79,9 @@
:on-change #(rf/dispatch [::pb/seek (js/parseInt (.. % -target -value) 10)])}]
[:div.readout
[:span (str "frame " frame " / " frames)]
[:span (str fps " fps")]
[:span (str "exposure " expose " → holds " (clock/exposed-frame frame expose))]
[:span (str "source " fps " fps")]
[:span (str "picture " picture-fps " fps")]
[:span (str "pose " (clock/picture-frame frame fps picture-fps expose))]
[:span (str (js/Math.round (* 100 rate)) "%")]
;; Measured in the loop, not derived from the clock: the whole question
;; while profiling is whether the painting keeps up with the clock, so a
@ -88,6 +91,19 @@
(str (.toFixed (or fps 0) 1) " paint/s"
(when (and drop (pos? drop))
(str " · " (.toFixed drop 2) " frames/paint")))])]
(when (= id current)
[:div.picture-rate
[:span "picture fps "]
(doall
(for [r (distinct (filter #(<= % fps) [8 12 15 24 fps]))]
^{:key r}
[:button {:class (when (= r picture-fps) "on")
:on-click #(rf/dispatch [::pb/set-picture-fps r])}
(if (= r fps) "source" (str r))]))])
[:label.source-path "source manifest "
[:input {:type "text" :value manifest-path :disabled loading?
:on-change #(rf/dispatch [::footage/set-manifest-path
(.. % -target -value)])}]]
(when status [:div.load-status status])]))
(defn- stage []