Preserve source cadence and sample picture fps after analysis
This commit is contained in:
parent
32683efccf
commit
06cf02db83
20 changed files with 303 additions and 87 deletions
2
.gitignore
vendored
2
.gitignore
vendored
|
|
@ -1,4 +1,6 @@
|
|||
frames/
|
||||
# local extracted takes for comparing source cadences
|
||||
/scratch/
|
||||
*.task
|
||||
*.take
|
||||
|
||||
|
|
|
|||
32
README.md
32
README.md
|
|
@ -42,7 +42,7 @@ wasm, which is fetched from a CDN on first use.
|
|||
For real footage:
|
||||
|
||||
```sh
|
||||
./extract.sh /path/to/clip.mov 12 # -> frames/*.png, audio.wav, manifest.json
|
||||
./extract.sh /path/to/clip.mov # -> frames/*.png, audio.wav, manifest.json
|
||||
```
|
||||
|
||||
then **Load frames**. MediaPipe's wasm is fetched from jsdelivr on first use;
|
||||
|
|
@ -52,16 +52,16 @@ Frames are pre-extracted rather than decoded in the page because browser video
|
|||
seeking is approximate and `requestVideoFrameCallback` only delivers frames at
|
||||
playback speed — neither gives a deterministic per-frame pass.
|
||||
|
||||
`manifest.json` records the true extraction rate. The page reads it rather than
|
||||
assuming, because a guessed fps desynchronises audio from picture — and sync is
|
||||
the one thing this view exists to show.
|
||||
`manifest.json` records the source rate. The extractor keeps every source frame;
|
||||
the page reads that rate because a guessed fps desynchronises audio from picture.
|
||||
Choosing a lower picture rate happens after analysis.
|
||||
|
||||
**Exposure** decides how often the picture gets a new drawing: rip at 24 and
|
||||
render `on 2s` for 12, `on 3s` for 8. The dense track and the audio are
|
||||
untouched, so it is a dropdown rather than a re-rip, and the export emits keys
|
||||
only on the grid instead of the same pose twice. Everything rides the same grid
|
||||
— mouth, eyes, teeth, plate — because a head cutting on the odd frames while the
|
||||
mouth cuts on the even ones reads as two performances laid over each other.
|
||||
**Picture fps** decides how often the finished roto gets a new pose. Analyze all
|
||||
source frames, then sample those frozen poses at 12, 24 or the source rate while
|
||||
keeping the original duration and audio. **Exposure** can hold a drawing across
|
||||
more than one picture slot. The tracing editor chooses source frames for cel
|
||||
references separately. Shared timing is the useful default for mouth, eyes,
|
||||
teeth and plate so their changes read as one performance.
|
||||
|
||||
**Audio is the playback clock**: `frame = floor(audio.currentTime * fps)`. A slow
|
||||
render loop therefore drops frames instead of drifting, and ½x / ¼x work by
|
||||
|
|
@ -322,13 +322,13 @@ the tool a person made by hand; everything else regenerates. They are not in the
|
|||
## Two kinds of sparseness
|
||||
|
||||
Sparseness has two unrelated causes, and conflating them was the original design
|
||||
error here. **Aesthetic** sparseness is set by the extraction rate — pick 12fps and
|
||||
you have already chosen your timing. **Labour** sparseness is a human drawing
|
||||
each one, and it binds only on the plate.
|
||||
error here. **Aesthetic** sparseness is chosen from the full analyzed source
|
||||
track at rendering time. **Labour** sparseness is a human drawing each cel and
|
||||
selecting which source frames to use as tracing references.
|
||||
|
||||
Aesthetic sparseness is the **exposure** control, not the extraction rate —
|
||||
making it a render-time grid means auditioning 12 against 24 costs a dropdown
|
||||
instead of a re-rip and a full re-detection.
|
||||
Aesthetic sparseness is the **picture fps** control, with exposure available for
|
||||
longer holds. Both happen after analysis, so auditioning 12 against 24 needs no
|
||||
re-extraction or re-detection.
|
||||
|
||||
So the mouth keeps **every** frame: it is traced, and therefore free. In limited
|
||||
animation lip sync is routinely the densest element, on 1s, while heads hold on
|
||||
|
|
|
|||
|
|
@ -333,11 +333,15 @@ Every node may map the frame it is evaluated at:
|
|||
```clojure
|
||||
:time {:mode :inherit} ; the default, and almost always right
|
||||
:time {:mode :map :expose 2 :offset -1 :rate 1.0 :loop? false}
|
||||
:time {:mode :map :source-fps 30 :sample-fps 12} ; root: lower picture cadence
|
||||
```
|
||||
|
||||
Three features that look unrelated are this one mechanism:
|
||||
|
||||
- **exposure** is `⌊f/n⌋·n`,
|
||||
- **picture fps** quantises source time to a chosen picture grid, then reads the
|
||||
latest source pose at or before that time; source analysis and audio keep their
|
||||
original cadence,
|
||||
- **mouth lead** is `f + k`,
|
||||
- **a symbol instance's timing** is `(f - at)·rate + in`, with optional looping.
|
||||
|
||||
|
|
@ -463,8 +467,11 @@ world transform of the node it rides:
|
|||
|
||||
Composed with image-pixels-to-local — **both axes divided by `imgH`**, never by
|
||||
their own dimension — the photo is registered with the shapes by construction,
|
||||
and an unregistered underlay is merely decorative. Which frame it shows, the
|
||||
current one or the held plate frame, is a UI choice and not a stored one.
|
||||
and an unregistered underlay is merely decorative. The tracing editor chooses
|
||||
which source frame to show under a cel. That reference choice is independent of
|
||||
the finished picture fps and does not change the dense analysis track. A cel can
|
||||
therefore use any useful source frame as its drawing reference, even when that
|
||||
frame is not one of the displayed picture poses.
|
||||
|
||||
A photo that has to sit *between* two drawn layers is the case that would make it
|
||||
a `:bitmap` node with an op of its own. Nothing wants that yet: a reference is
|
||||
|
|
@ -527,6 +534,7 @@ Proof that it covers what exists, not just what is wanted:
|
|||
| painted background cel | node per layer, `[:geom :pts]` **framed**, `[:style :color]` framed |
|
||||
| `mouth lead` | `:time {:offset k}` on performance nodes only |
|
||||
| `exposure` | `:time {:expose n}` on the clip root, inherited |
|
||||
| picture fps | `:time {:source-fps s :sample-fps p}` on the clip root, applied after analysis |
|
||||
| hand correction | an `:over` layer, `:offset` or `:replace` |
|
||||
|
||||
The brow row is the one worth looking at twice. `docs/design.md` argues at length
|
||||
|
|
|
|||
|
|
@ -49,6 +49,11 @@ overrides and kept-frame sets are all in clip-frame space, so a clip slides on
|
|||
the timeline without a single stored number changing. Exposure and lead are
|
||||
transforms *within* clip space:
|
||||
|
||||
Detection retains every source frame. A chosen picture fps samples the frozen
|
||||
roto in clip time; it changes neither source-frame count nor the audio clock.
|
||||
The set of source frames an artist uses as cel tracing references is another
|
||||
selection, independent of the picture fps.
|
||||
|
||||
```clojure
|
||||
(defn pose-frame [clip cf]
|
||||
(-> cf (expose (:exposure clip)) (shift (:lead clip) (count-frames clip))))
|
||||
|
|
@ -73,7 +78,7 @@ project
|
|||
│ ├── scene node tree
|
||||
│ ├── channels node+property -> keyframe stream, in cf
|
||||
│ ├── overrides (node, property, cf) -> value, applied last
|
||||
│ └── cels painted vector layers, per exposure slot
|
||||
│ └── cels painted vector layers, keyed by chosen clip frames
|
||||
└── sequence[] clip placements: {clip-id, at, in, out}
|
||||
```
|
||||
|
||||
|
|
@ -81,6 +86,10 @@ project
|
|||
it a name is most of the work: `state` in `app.js` is a clip with its analysis
|
||||
inlined and its palette global.
|
||||
|
||||
Cel keys select where drawings begin and how long they hold. The source frames
|
||||
shown beneath a cel while tracing are chosen independently, and picture fps
|
||||
only controls which analyzed pose the finished roto displays at a given time.
|
||||
|
||||
### Two things called "track"
|
||||
|
||||
`docs/design.md` says "dense track" for the landmark stream. A timeline also
|
||||
|
|
|
|||
|
|
@ -4,8 +4,10 @@ Self-contained. You should not need any prior conversation to execute this.
|
|||
|
||||
**Implementation status (2026-09-27):** steps 0–6 are in the CLJS frontend.
|
||||
Step 6 reads extracted footage from the manifest, detects landmarks with local
|
||||
MediaPipe assets, and runs the same freeze path as the synthetic take. Step 7 is
|
||||
next: eyes, brows and the pixel-derived mouth interior.
|
||||
MediaPipe assets at full source cadence, and runs the same freeze path as the
|
||||
synthetic take. The scene time map can sample the frozen roto at a lower picture
|
||||
fps without changing source analysis, duration or audio. Step 7 is next: eyes,
|
||||
brows and the pixel-derived mouth interior.
|
||||
|
||||
## What arthur is
|
||||
|
||||
|
|
@ -302,6 +304,10 @@ MediaPipe interop behind one namespace; real frames, real audio, real fps from
|
|||
the manifest. **Vendor the wasm** rather than fetching from jsdelivr — it is
|
||||
currently the only thing in the tool that silently requires a network.
|
||||
|
||||
Decode every source frame for analysis. A lower output picture fps is a time map
|
||||
over frozen channels, not a reduced detection track. Selecting source frames to
|
||||
trace into cels is independent again and remains outside this port's paint scope.
|
||||
|
||||
**Done:** real footage plays back as a rotoscoped mouth.
|
||||
|
||||
### 7 — the rest of measure
|
||||
|
|
|
|||
93
extract.sh
93
extract.sh
|
|
@ -7,32 +7,77 @@
|
|||
# pass. A PNG sequence is exact, instantly seekable, and reproducible.
|
||||
#
|
||||
# Audio comes out alongside because the page uses it as the PLAYBACK CLOCK -
|
||||
# frame = floor(audio.currentTime * fps) - so picture and sound cannot drift
|
||||
# apart no matter how long the shot is or how slow the render loop runs.
|
||||
# frame = floor(audio.currentTime * fps). Detection sees every decoded source
|
||||
# frame; a lower drawing rate is a later playback choice, never an extraction
|
||||
# choice. This script accepts CFR footage because frame-index timing needs a
|
||||
# single rate. VFR needs per-frame timestamps in the manifest first.
|
||||
set -euo pipefail
|
||||
|
||||
src="${1:?usage: ./extract.sh CLIP [FPS] [OUTDIR]}"
|
||||
fps="${2:-12}"
|
||||
out="${3:-frames}"
|
||||
|
||||
rm -rf "$out"; mkdir -p "$out"
|
||||
ffmpeg -hide_banner -loglevel warning -i "$src" -vf "fps=$fps" "$out/%04d.png"
|
||||
count=$(ls -1 "$out" | wc -l)
|
||||
|
||||
# Mono is enough for judging sync and halves the file. Absent audio is not fatal.
|
||||
if ffprobe -v error -select_streams a:0 -show_entries stream=codec_type \
|
||||
-of csv=p=0 "$src" 2>/dev/null | grep -q audio; then
|
||||
ffmpeg -hide_banner -loglevel warning -y -i "$src" -vn -ac 1 -ar 44100 audio.wav
|
||||
audio='"audio.wav"'
|
||||
echo "audio -> audio.wav"
|
||||
else
|
||||
audio='null'
|
||||
echo "no audio stream"
|
||||
src="${1:?usage: ./extract.sh CLIP [BUNDLE_DIR]}"
|
||||
bundle="${2:-.}"
|
||||
if [[ "$bundle" =~ ^[0-9]+([.][0-9]+)?$ ]]; then
|
||||
echo "The FPS argument was removed: extraction always keeps the source rate. Use a directory as argument 2." >&2
|
||||
exit 2
|
||||
fi
|
||||
if [[ "$bundle" = /* || "$bundle" = *..* ]]; then
|
||||
echo "BUNDLE_DIR must be a relative directory inside this repo" >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
# The page must know the true extraction rate: if it guessed, audio and picture
|
||||
# would drift. Source of truth lives here, next to the frames it describes.
|
||||
printf '{"fps":%s,"frames":%s,"dir":"%s","audio":%s,"source":"%s"}\n' \
|
||||
"$fps" "$count" "$out" "$audio" "$(basename "$src")" > manifest.json
|
||||
probe=$(ffprobe -v error -select_streams v:0 \
|
||||
-show_entries stream=r_frame_rate,avg_frame_rate,nb_frames \
|
||||
-of json "$src")
|
||||
fps=$(python3 -c '
|
||||
import json, sys
|
||||
from fractions import Fraction
|
||||
streams = json.load(sys.stdin).get("streams", [])
|
||||
if not streams:
|
||||
raise SystemExit("no video stream in source")
|
||||
s = streams[0]
|
||||
nominal = Fraction(s["r_frame_rate"])
|
||||
average = Fraction(s["avg_frame_rate"])
|
||||
if nominal <= 0 or average <= 0 or abs(float(nominal / average) - 1) > 0.001:
|
||||
raise SystemExit("variable-frame-rate source needs timestamp-aware playback; refusing to guess its fps")
|
||||
print(float(average))
|
||||
' <<< "$probe")
|
||||
|
||||
echo "$count frames at ${fps}fps -> $out/ (manifest.json written)"
|
||||
if [[ "$bundle" = "." ]]; then
|
||||
dir="frames"; audio_path="audio.wav"; manifest_path="manifest.json"
|
||||
else
|
||||
dir="${bundle%/}/frames"
|
||||
audio_path="${bundle%/}/audio.wav"
|
||||
manifest_path="${bundle%/}/manifest.json"
|
||||
fi
|
||||
|
||||
rm -rf "$dir"; mkdir -p "$dir"
|
||||
ffmpeg -hide_banner -loglevel warning -i "$src" -fps_mode passthrough "$dir/%04d.png"
|
||||
count=$(find "$dir" -maxdepth 1 -name '*.png' -type f | wc -l | tr -d ' ')
|
||||
|
||||
expected=$(python3 -c 'import json,sys; print(json.load(sys.stdin)["streams"][0].get("nb_frames", ""))' <<< "$probe")
|
||||
if [[ "$expected" =~ ^[0-9]+$ && "$count" != "$expected" ]]; then
|
||||
echo "decoded $count frames but source reports $expected; refusing an inaccurate manifest" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Mono is enough for judging sync and halves the file. A silent clock lets a
|
||||
# mute source use the same audio-driven transport.
|
||||
if ffprobe -v error -select_streams a:0 -show_entries stream=codec_type \
|
||||
-of csv=p=0 "$src" 2>/dev/null | grep -q audio; then
|
||||
ffmpeg -hide_banner -loglevel warning -y -i "$src" -vn -ac 1 -ar 44100 "$audio_path"
|
||||
else
|
||||
duration=$(python3 -c 'import sys; print(int(sys.argv[1]) / float(sys.argv[2]))' "$count" "$fps")
|
||||
ffmpeg -hide_banner -loglevel warning -y -f lavfi -i anullsrc=r=44100:cl=mono \
|
||||
-t "$duration" -c:a pcm_s16le "$audio_path"
|
||||
fi
|
||||
|
||||
# JSON escaping belongs to a JSON writer, especially for source filenames.
|
||||
python3 - "$fps" "$count" "$dir" "$audio_path" "$src" "$manifest_path" <<'PY'
|
||||
import json, os, sys
|
||||
fps, count, frames, audio, source, path = sys.argv[1:]
|
||||
with open(path, "w") as out:
|
||||
json.dump({"fps": float(fps), "frames": int(count), "dir": frames,
|
||||
"audio": audio, "source": os.path.basename(source)}, out)
|
||||
out.write("\n")
|
||||
PY
|
||||
|
||||
echo "$count source frames at ${fps}fps -> $dir/ ($manifest_path written)"
|
||||
|
|
|
|||
|
|
@ -87,16 +87,32 @@ and real footage use `src/arthur/flow/take.cljs` for the measurement order and
|
|||
From the repo root, extract a clip, then click **load frames** in the CLJS app:
|
||||
|
||||
```sh
|
||||
./extract.sh /path/to/clip.mov 12
|
||||
./extract.sh /path/to/clip.mov
|
||||
```
|
||||
|
||||
This writes `frames/0001.png` onward, `audio.wav`, and `manifest.json` at the
|
||||
repo root. The manifest supplies the exact frame count, fps and audio path.
|
||||
This keeps every source frame and writes `frames/0001.png` onward, `audio.wav`,
|
||||
and `manifest.json` at the repo root. The manifest supplies the exact frame
|
||||
count, source fps and audio path. Variable frame rate sources are rejected until
|
||||
the manifest and clock carry per-frame timestamps.
|
||||
|
||||
To keep multiple takes or compare with a previous extraction, pass a bundle
|
||||
directory and enter its manifest path in the app:
|
||||
|
||||
```sh
|
||||
./extract.sh /path/to/clip.mov scratch/my-take
|
||||
# source manifest field: /scratch/my-take/manifest.json
|
||||
```
|
||||
|
||||
`scratch/` is ignored by Git. The directory contains its own frames, audio and
|
||||
manifest, so extracting it does not replace another take's files.
|
||||
Loading detects one face per frame, freezes the measured mouth into channels,
|
||||
and adds a button for the footage clip. Detection happens once when you load;
|
||||
playback only resolves channels and paints. Frames without a detection remain
|
||||
marked absent even though their neighbouring poses are used to condition the
|
||||
track. The stage stays 320×200 regardless of the footage dimensions.
|
||||
track. The stage stays 320×200 regardless of the footage dimensions. Real
|
||||
footage starts at the source picture rate. The **picture fps** buttons sample the
|
||||
frozen roto at lower rates while the source track, duration and audio clock stay
|
||||
unchanged. Picking frames to trace into cels is a separate future editing step.
|
||||
|
||||
MediaPipe's JS, wasm and model are under `public/mediapipe/` and served locally.
|
||||
No CDN is used by this app. See that directory's README for provenance.
|
||||
|
|
|
|||
|
|
@ -31,6 +31,12 @@
|
|||
.scrub { width: 100%; margin: 10px 0 6px; }
|
||||
.readout { display: flex; gap: 18px; opacity: .55; font-size: 12px; }
|
||||
.readout .warn { color: #d98f5a; opacity: 1; }
|
||||
.picture-rate { display: flex; align-items: center; gap: 6px; margin-top: 7px;
|
||||
font-size: 12px; }
|
||||
.source-path { display: block; margin-top: 8px; font-size: 12px; opacity: .7; }
|
||||
.source-path input { width: 300px; margin-left: 8px; padding: 3px 5px;
|
||||
color: var(--fg); background: #1c1f2b; border: 1px solid #2b3040;
|
||||
font: inherit; }
|
||||
.load-status { margin-top: 6px; font-size: 12px; opacity: .75; }
|
||||
.note { opacity: .35; font-size: 12px; max-width: 640px; }
|
||||
</style>
|
||||
|
|
|
|||
|
|
@ -97,3 +97,8 @@
|
|||
rather than only inside the scene."
|
||||
[f expose]
|
||||
(node/expose f expose))
|
||||
|
||||
(defn picture-frame
|
||||
"The source pose displayed at f after picture-rate sampling and exposure."
|
||||
[f source-fps picture-fps expose]
|
||||
(node/expose (node/sample-frame f source-fps picture-fps) expose))
|
||||
|
|
|
|||
|
|
@ -23,7 +23,8 @@
|
|||
because there is one clip per scene today. Copying them by hand into this table
|
||||
is how one of them comes to disagree with the scene it describes."
|
||||
[label scene store]
|
||||
(merge {:label label :scene scene :store store :audio "/audio.wav"}
|
||||
(merge {:label label :scene scene :store store :audio "/audio.wav"
|
||||
:display-fps (:fps scene)}
|
||||
(select-keys scene [:fps :frames :width :height])))
|
||||
|
||||
(def scenes
|
||||
|
|
@ -50,9 +51,10 @@
|
|||
;; footage's. That is what deleting `makeXform` buys — the framing became a
|
||||
;; transform on a node, so nothing downstream of the freeze knows the frame
|
||||
;; size — and it is why ui/player no longer hardcodes 320x200.
|
||||
:clip (select-keys (:take scenes) [:fps :frames :width :height :audio])
|
||||
:clip (select-keys (:take scenes) [:fps :frames :width :height :audio :display-fps])
|
||||
|
||||
:footage {:id nil :label nil :loading? false :status nil}
|
||||
:footage {:id nil :label nil :loading? false :status nil
|
||||
:manifest-path "/manifest.json"}
|
||||
|
||||
;; --- transport ---
|
||||
;;
|
||||
|
|
|
|||
|
|
@ -85,6 +85,22 @@
|
|||
[f n]
|
||||
(if (and n (> n 1)) (* (js/Math.floor (/ f n)) n) f))
|
||||
|
||||
(defn sample-frame
|
||||
"Pick a source frame for a lower picture rate without changing clip time.
|
||||
|
||||
The input is already an integer source frame from the audio clock. Its time is
|
||||
f/source-fps. Quantise that time to the picture grid, then read the latest
|
||||
source frame at or before it. The result is always an integer and never from
|
||||
the future, including when the rates do not divide (30 source → 24 picture)."
|
||||
[f source-fps picture-fps]
|
||||
(if (and source-fps picture-fps
|
||||
(pos? source-fps) (pos? picture-fps)
|
||||
(< picture-fps source-fps))
|
||||
(min f (js/Math.floor
|
||||
(* (js/Math.floor (/ (* f picture-fps) source-fps))
|
||||
(/ source-fps picture-fps))))
|
||||
f))
|
||||
|
||||
(defn local-frame
|
||||
"Apply a node's time map to the frame it was handed by its parent.
|
||||
|
||||
|
|
@ -99,7 +115,8 @@
|
|||
as two performances; offset is PER-NODE by design, because mouth lead applies
|
||||
to performance nodes and not to the plate, which is the entire point of it."
|
||||
[n f]
|
||||
(let [{:keys [mode offset rate] ex :expose :or {mode :inherit}} (:time n)]
|
||||
(let [{:keys [mode offset rate source-fps sample-fps]
|
||||
ex :expose :or {mode :inherit}} (:time n)]
|
||||
(if (= mode :inherit)
|
||||
f
|
||||
(do
|
||||
|
|
@ -110,7 +127,11 @@
|
|||
(when (and rate (not= rate 1.0) (not= rate 1))
|
||||
(throw (ex-info "time map :rate is symbol timing and symbols are not built (port-plan step 2 scope)"
|
||||
{:node (:id n) :time (:time n)})))
|
||||
(when (and sample-fps (not (and source-fps (pos? source-fps))))
|
||||
(throw (ex-info "picture sampling needs a positive source fps"
|
||||
{:node (:id n) :time (:time n)})))
|
||||
(cond-> f
|
||||
sample-fps (sample-frame source-fps sample-fps)
|
||||
ex (expose ex)
|
||||
offset (+ offset))))))
|
||||
|
||||
|
|
@ -256,4 +277,3 @@
|
|||
(into (for [[path c] (:channels n)
|
||||
p (ch/problems c)]
|
||||
(str "channel " (pr-str path) ": " p))))))
|
||||
|
||||
|
|
|
|||
|
|
@ -43,13 +43,11 @@
|
|||
|
||||
(defn- build-clip [manifest {:keys [dense detected dimensions missing first-real]}]
|
||||
(let [[w h] dimensions
|
||||
params (merge take/knobs
|
||||
{:name "footage" :fps (:fps manifest) :aspect (/ w h)
|
||||
:stage [320 200] :expose 2 :head :as-filmed
|
||||
:analysis (str "mediapipe:1.0.1/" (:source manifest))})
|
||||
frozen (take/build params {:dense dense :detected detected})
|
||||
frozen (take/footage manifest {:dense dense :detected detected
|
||||
:dimensions dimensions})
|
||||
scene (:scene frozen)]
|
||||
(assoc (select-keys scene [:fps :frames :width :height])
|
||||
:display-fps (:fps scene)
|
||||
:scene scene :store (:store frozen)
|
||||
;; Re-extraction often overwrites audio.wav under the same name. A new
|
||||
;; URL makes the element fetch the new sound when this clip is loaded.
|
||||
|
|
@ -64,8 +62,8 @@
|
|||
|
||||
(rf/reg-fx
|
||||
::begin!
|
||||
(fn [_]
|
||||
(-> (ingest/manifest!)
|
||||
(fn [path]
|
||||
(-> (ingest/manifest! path)
|
||||
(.then (fn [manifest]
|
||||
(rf/dispatch [::progress "loading MediaPipe…"])
|
||||
(-> (detect/landmarker!)
|
||||
|
|
@ -88,7 +86,11 @@
|
|||
{:db (assoc db :footage (assoc (:footage db) :loading? true
|
||||
:status "reading manifest.json…"))
|
||||
::pb/pause! nil
|
||||
::begin! nil})))
|
||||
::begin! (get-in db [:footage :manifest-path])})))
|
||||
|
||||
(rf/reg-event-db
|
||||
::set-manifest-path
|
||||
(fn [db [_ path]] (assoc-in db [:footage :manifest-path] path)))
|
||||
|
||||
(rf/reg-event-db
|
||||
::progress
|
||||
|
|
@ -106,9 +108,9 @@
|
|||
(let [clip (store/entry id)]
|
||||
{:db (-> db
|
||||
(assoc :scene/current id
|
||||
:clip (select-keys clip [:fps :frames :width :height :audio])
|
||||
:footage {:id id :label (:label clip)
|
||||
:loading? false :status summary})
|
||||
:clip (select-keys clip [:fps :frames :width :height :audio :display-fps])
|
||||
:footage (assoc (:footage db) :id id :label (:label clip)
|
||||
:loading? false :status summary))
|
||||
(assoc-in [:playback :frame] 0)
|
||||
(assoc-in [:playback :playing?] false))
|
||||
::pb/pause! nil})))
|
||||
|
|
|
|||
|
|
@ -65,6 +65,13 @@
|
|||
{:db (assoc-in db [:playback :rate] r)
|
||||
::rate! r}))
|
||||
|
||||
(rf/reg-event-db
|
||||
::set-picture-fps
|
||||
(fn [db [_ target]]
|
||||
(if (and (number? target) (pos? target) (<= target (fps db)))
|
||||
(assoc-in db [:clip :display-fps] target)
|
||||
db)))
|
||||
|
||||
;; --- effects: every DOM touch on the audio element is one of these ---
|
||||
|
||||
(rf/reg-fx ::play! (fn [_] (clock/play!)))
|
||||
|
|
@ -98,7 +105,7 @@
|
|||
;; The stage travels with the clip: two clips may be different
|
||||
;; sizes, and the raster the loop paints into is the clip's, not
|
||||
;; the app's.
|
||||
(assoc :clip (select-keys clip [:fps :frames :width :height :audio]))
|
||||
(assoc :clip (select-keys clip [:fps :frames :width :height :audio :display-fps]))
|
||||
(assoc-in [:playback :frame] 0)
|
||||
(assoc-in [:playback :playing?] false))
|
||||
::pause! nil
|
||||
|
|
|
|||
|
|
@ -275,6 +275,31 @@
|
|||
;; ---------------------------------------------------------------------------
|
||||
;; the face, onto the stage
|
||||
|
||||
(defn- motion-placement
|
||||
"An editable default framing for a real take. Fit the observed eye/nose and
|
||||
mouth motion within the stage; otherwise a downward head move can push the
|
||||
mouth below a fixed 320×200 canvas even while it stays in the source image."
|
||||
[[w h] {:keys [rigid transforms outer detected]}]
|
||||
(let [points (mapcat (fn [i]
|
||||
(when (or (nil? detected) (nth detected i))
|
||||
(concat (nth rigid i)
|
||||
(geom/apply-sim-all (invert (nth transforms i))
|
||||
(nth outer i)))))
|
||||
(range (count outer)))
|
||||
xs (map :x points)
|
||||
ys (map :y points)
|
||||
x0 (reduce min xs)
|
||||
x1 (reduce max xs)
|
||||
y0 (reduce min ys)
|
||||
y1 (reduce max ys)
|
||||
cx (/ (+ x0 x1) 2)
|
||||
cy (/ (+ y0 y1) 2)
|
||||
k (min (/ (* 0.8 w) (max 1e-9 (- x1 x0)))
|
||||
(/ (* 0.8 h) (max 1e-9 (- y1 y0))))]
|
||||
{[:xform :anchor] (ch/framed [cx cy])
|
||||
[:xform :scale] (ch/framed [k k])
|
||||
[:xform :pos] (ch/framed [(- (/ w 2) cx) (- (/ h 2) cy)])}))
|
||||
|
||||
(defn face-placement
|
||||
"The face's transform on the stage, as FRAMED channels.
|
||||
|
||||
|
|
@ -284,9 +309,14 @@
|
|||
ordinary node. `makeXform` made the same decision and then baked it into every
|
||||
vertex, where nothing could ever revise it.
|
||||
|
||||
Derived from the reference rigid configuration, which is what the freeze has:
|
||||
the face oval is not measured, because its only consumers in the prototype were
|
||||
that transform and the placeholder plate outline. Two numbers:
|
||||
The synthetic take uses the reference rigid configuration for its default.
|
||||
Real footage can request `:fit-motion?`: its default fits the observed rigid
|
||||
and lip motion in the stage so a head move does not send the mouth off canvas.
|
||||
Both are ordinary editable transforms on :face, never baked into the geometry.
|
||||
The face oval is not measured, because its only consumers in the prototype were
|
||||
the old baked framing transform and the placeholder plate outline.
|
||||
|
||||
For the synthetic default, two numbers:
|
||||
|
||||
SCALE is stage pixels per image height, set so the reference's eye-corner span
|
||||
is 40% of the stage width. Landmark-free — it is the rigid configuration's own
|
||||
|
|
@ -304,7 +334,9 @@
|
|||
landmarks are eyes and nose — the upper middle of a face — so a quarter down
|
||||
leaves the jaw and the mouth on the stage. Whatever hangs off is clipped, which
|
||||
is not a feature to add: every fill in `domain/raster` clamps already."
|
||||
[{:keys [stage]} {:keys [ref]}]
|
||||
[{:keys [stage fit-motion?]} {:keys [ref] :as inputs}]
|
||||
(if fit-motion?
|
||||
(motion-placement stage inputs)
|
||||
(let [[w h] stage
|
||||
c (geom/centroid ref)
|
||||
span (- (reduce max (map :x ref)) (reduce min (map :x ref)))
|
||||
|
|
@ -312,7 +344,7 @@
|
|||
{[:xform :anchor] (ch/framed [(:x c) (:y c)])
|
||||
[:xform :scale] (ch/framed [k k])
|
||||
[:xform :pos] (ch/framed [(- (/ w 2) (:x c))
|
||||
(- (* 0.25 h) (:y c))])}))
|
||||
(- (* 0.25 h) (:y c))])})))
|
||||
|
||||
;; ---------------------------------------------------------------------------
|
||||
;; the clip
|
||||
|
|
@ -327,6 +359,8 @@
|
|||
`(count outer)`, because a freeze that could disagree with its
|
||||
own input about the length of the take would.
|
||||
:stage [w h] project dimensions, INDEPENDENT of the footage
|
||||
:fit-motion? choose an editable real-footage default that keeps the
|
||||
observed feature motion on stage
|
||||
:expose the clip root's exposure grid, inherited by everything
|
||||
:verts the lip rings' vertex budget
|
||||
:aperture-cut fraction of the take's peak aperture below which the mouth
|
||||
|
|
|
|||
|
|
@ -13,8 +13,8 @@
|
|||
(assoc m :fps fps :frames frames)))
|
||||
|
||||
(defn manifest!
|
||||
[]
|
||||
(-> (js/fetch "/manifest.json" #js {:cache "no-store"})
|
||||
[path]
|
||||
(-> (js/fetch (str "/" (str/replace path #"^/+" "")) #js {:cache "no-store"})
|
||||
(.then (fn [response]
|
||||
(when-not (.-ok response)
|
||||
(throw (ex-info "manifest.json was not found; run extract.sh first"
|
||||
|
|
|
|||
|
|
@ -24,3 +24,16 @@
|
|||
(defn build
|
||||
[params inputs]
|
||||
(freeze/clip params (measure params inputs)))
|
||||
|
||||
(defn footage
|
||||
"A real manifest and its detected landmarks through the same measurement and
|
||||
freeze path as the synthetic take. The source cadence stays in :fps; picture
|
||||
sampling is a root time map applied only after this artifact exists."
|
||||
[manifest {:keys [dense detected dimensions]}]
|
||||
(let [[w h] dimensions
|
||||
params (merge knobs
|
||||
{:name "footage" :fps (:fps manifest) :aspect (/ w h)
|
||||
:stage [320 200] :fit-motion? true
|
||||
:expose 1 :head :as-filmed
|
||||
:analysis (str "mediapipe:1.0.1/" (:source manifest))})]
|
||||
(build params {:dense dense :detected detected})))
|
||||
|
|
|
|||
|
|
@ -10,6 +10,7 @@
|
|||
(rf/reg-sub ::loop? (fn [db _] (get-in db [:playback :loop?])))
|
||||
(rf/reg-sub ::muted? (fn [db _] (get-in db [:playback :muted?])))
|
||||
(rf/reg-sub ::fps (fn [db _] (get-in db [:clip :fps])))
|
||||
(rf/reg-sub ::display-fps (fn [db _] (get-in db [:clip :display-fps])))
|
||||
(rf/reg-sub ::frames (fn [db _] (get-in db [:clip :frames])))
|
||||
;; The stage, in pixels. On the clip because project dimensions are independent
|
||||
;; of the footage — see flow/freeze/face-placement — so the canvas and the raster
|
||||
|
|
|
|||
|
|
@ -10,11 +10,28 @@
|
|||
(:require [arthur.domain.palette :as pal]
|
||||
[arthur.domain.scene :as scene]
|
||||
[arthur.footage.store :as footage]
|
||||
[arthur.subs.playback :as playback]
|
||||
[re-frame.core :as rf]))
|
||||
|
||||
(rf/reg-sub ::scene-id (fn [db _] (:scene/current db)))
|
||||
|
||||
(rf/reg-sub ::scene (fn [db _] (:scene (footage/entry (:scene/current db)))))
|
||||
(rf/reg-sub
|
||||
::base-scene
|
||||
:<- [::scene-id]
|
||||
(fn [id _] (:scene (footage/entry id))))
|
||||
|
||||
(rf/reg-sub
|
||||
::scene
|
||||
:<- [::base-scene]
|
||||
:<- [::playback/display-fps]
|
||||
(fn [[scene picture-fps] _]
|
||||
;; This changes only the scene's root time map. The dense source track stays
|
||||
;; at its native rate, and the audio clock still advances through source time.
|
||||
(if (and scene picture-fps (< picture-fps (:fps scene)))
|
||||
(-> scene
|
||||
(assoc-in [:nodes :root :time :source-fps] (:fps scene))
|
||||
(assoc-in [:nodes :root :time :sample-fps] picture-fps))
|
||||
scene)))
|
||||
|
||||
(rf/reg-sub
|
||||
::exposure
|
||||
|
|
@ -45,11 +62,12 @@
|
|||
|
||||
(rf/reg-sub
|
||||
::store
|
||||
(fn [db _]
|
||||
:<- [::scene-id]
|
||||
(fn [id _]
|
||||
;; Tier 2, behind a handle, and never in app-db itself — what is in the db is
|
||||
;; the id of the clip whose blocks these are. The hand-written demo has none;
|
||||
;; the swarm is entirely dense.
|
||||
(:store (footage/entry (:scene/current db)))))
|
||||
(:store (footage/entry id))))
|
||||
|
||||
(rf/reg-sub
|
||||
::resolver
|
||||
|
|
|
|||
|
|
@ -167,12 +167,18 @@
|
|||
;; the next one lands where the audio already is, so the failure mode is
|
||||
;; a visible stutter rather than an invisible slide out of sync.
|
||||
f (if live? (clock/frame fps frames) frame)]
|
||||
;; A paused scrub is an explicit seek, not a playback sample. Clear the
|
||||
;; meter before the next play so its first frame cannot count the seek as
|
||||
;; dropped frames or leave a stale low rate in the transport.
|
||||
(when (and (not live?) (:t0 @meter-state))
|
||||
(reset! meter-state {:t0 nil :paints 0 :samples 0 :advanced 0 :prev nil})
|
||||
(reset! meter {:fps 0 :drop 0}))
|
||||
;; `ready?` gates the bookkeeping as well as the draw: recording a frame as
|
||||
;; painted when it was not is how the canvas stays empty forever.
|
||||
(when (and ready? f (not= f (:last @state)))
|
||||
(swap! state assoc :last f)
|
||||
(paint! f)
|
||||
(meter! f)
|
||||
(when live? (meter! f))
|
||||
(when live? (rf/dispatch [::pb/tick f])))
|
||||
;; The audio ending is the authority on playback having stopped; nothing
|
||||
;; counts frames to notice it.
|
||||
|
|
|
|||
|
|
@ -35,8 +35,10 @@
|
|||
frame @(rf/subscribe [::sub/frame])
|
||||
frames @(rf/subscribe [::sub/frames])
|
||||
fps @(rf/subscribe [::sub/fps])
|
||||
picture-fps @(rf/subscribe [::sub/display-fps])
|
||||
current @(rf/subscribe [::render/scene-id])
|
||||
expose @(rf/subscribe [::render/exposure])
|
||||
{:keys [id label loading? status]} @(rf/subscribe [::sub/footage])]
|
||||
{:keys [id label loading? status manifest-path]} @(rf/subscribe [::sub/footage])]
|
||||
[:div.transport
|
||||
[:div.row
|
||||
[:button {:on-click #(rf/dispatch [::pb/toggle])}
|
||||
|
|
@ -77,8 +79,9 @@
|
|||
:on-change #(rf/dispatch [::pb/seek (js/parseInt (.. % -target -value) 10)])}]
|
||||
[:div.readout
|
||||
[:span (str "frame " frame " / " frames)]
|
||||
[:span (str fps " fps")]
|
||||
[:span (str "exposure " expose " → holds " (clock/exposed-frame frame expose))]
|
||||
[:span (str "source " fps " fps")]
|
||||
[:span (str "picture " picture-fps " fps")]
|
||||
[:span (str "pose " (clock/picture-frame frame fps picture-fps expose))]
|
||||
[:span (str (js/Math.round (* 100 rate)) "%")]
|
||||
;; Measured in the loop, not derived from the clock: the whole question
|
||||
;; while profiling is whether the painting keeps up with the clock, so a
|
||||
|
|
@ -88,6 +91,19 @@
|
|||
(str (.toFixed (or fps 0) 1) " paint/s"
|
||||
(when (and drop (pos? drop))
|
||||
(str " · " (.toFixed drop 2) " frames/paint")))])]
|
||||
(when (= id current)
|
||||
[:div.picture-rate
|
||||
[:span "picture fps "]
|
||||
(doall
|
||||
(for [r (distinct (filter #(<= % fps) [8 12 15 24 fps]))]
|
||||
^{:key r}
|
||||
[:button {:class (when (= r picture-fps) "on")
|
||||
:on-click #(rf/dispatch [::pb/set-picture-fps r])}
|
||||
(if (= r fps) "source" (str r))]))])
|
||||
[:label.source-path "source manifest "
|
||||
[:input {:type "text" :value manifest-path :disabled loading?
|
||||
:on-change #(rf/dispatch [::footage/set-manifest-path
|
||||
(.. % -target -value)])}]]
|
||||
(when status [:div.load-status status])]))
|
||||
|
||||
(defn- stage []
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue