2026-09-24 14:38:07 -04:00
|
|
|
|
<!doctype html>
|
|
|
|
|
|
<html lang="en">
|
|
|
|
|
|
<head>
|
|
|
|
|
|
<meta charset="utf-8">
|
|
|
|
|
|
<meta name="viewport" content="width=device-width, initial-scale=1">
|
Become arthur: a standalone suite, not an Animator Pro front-end
The test renderer turned out to be the product. Everything that decides how the
work looks - stabilisation, reduction, timing, frame removal, palette - already
happens here, and the flat indexed output already reads the way it should.
The reason to leave is in the original design's own rule: never make a timing
decision that requires a full render to evaluate. Honouring that moved every
judgement out of Animator Pro, which left the host doing nothing but writing a
file, in exchange for modal UI, minutes-long renders, one-level undo, FLX delta
invariants, a single tween state and a cel singleton.
What does NOT change is the constraint. 320x200, indexed palette, flat fills,
no antialiasing - inherited, but load-bearing rather than accidental. The
rasteriser writes palette indices and expands to RGBA only at the end precisely
so nothing can soften an edge. Modern conveniences belong in the workflow.
Adds docs/design.md: the principles, carried over without the Poco/FLX/cel
machinery, plus architecture and an honest list of what is missing - the
largest gap being that plates still have nowhere to be drawn.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:47:41 -04:00
|
|
|
|
<title>arthur</title>
|
2026-09-24 14:38:07 -04:00
|
|
|
|
<style>
|
2026-09-24 14:51:15 -04:00
|
|
|
|
:root { --bg:#0b0d13; --panel:#12151f; --line:#232836; --fg:#e8eaf0;
|
|
|
|
|
|
--dim:#8891a5; --accent:#fbbf24; --ok:#4ade80; --err:#f87171; }
|
2026-09-24 14:38:07 -04:00
|
|
|
|
* { box-sizing: border-box; }
|
2026-09-24 14:51:15 -04:00
|
|
|
|
body { margin:0; background:var(--bg); color:var(--fg);
|
|
|
|
|
|
font:13px/1.5 ui-monospace,SFMono-Regular,Menlo,monospace; }
|
|
|
|
|
|
header { padding:12px 18px; border-bottom:1px solid var(--line);
|
|
|
|
|
|
display:flex; gap:10px; align-items:center; flex-wrap:wrap; }
|
|
|
|
|
|
h1 { font-size:14px; margin:0 8px 0 0; letter-spacing:.04em; }
|
|
|
|
|
|
h1 span { color:var(--dim); font-weight:400; }
|
|
|
|
|
|
main { padding:16px 18px 40px; display:flex; flex-direction:column; gap:16px; }
|
|
|
|
|
|
.row { display:flex; gap:16px; flex-wrap:wrap; }
|
|
|
|
|
|
.panel { background:var(--panel); border:1px solid var(--line); border-radius:4px; padding:12px; }
|
|
|
|
|
|
.panel h2 { font-size:11px; text-transform:uppercase; letter-spacing:.08em;
|
|
|
|
|
|
color:var(--dim); margin:0 0 8px; font-weight:500; }
|
|
|
|
|
|
canvas { display:block; image-rendering:pixelated; max-width:100%; border-radius:2px; }
|
|
|
|
|
|
button,input[type=text],select { font:inherit; background:#1c2130; color:var(--fg);
|
|
|
|
|
|
border:1px solid var(--line); border-radius:3px; padding:5px 10px; cursor:pointer; }
|
|
|
|
|
|
button:hover { border-color:var(--accent); }
|
|
|
|
|
|
input[type=text] { cursor:text; }
|
|
|
|
|
|
label.ctl { display:grid; grid-template-columns:132px 1fr 46px; gap:10px;
|
|
|
|
|
|
align-items:center; margin-bottom:6px; }
|
|
|
|
|
|
label.ctl span:first-child { color:var(--dim); }
|
|
|
|
|
|
label.ctl output { text-align:right; color:var(--accent); }
|
|
|
|
|
|
input[type=range] { width:100%; accent-color:var(--accent); }
|
|
|
|
|
|
#status { color:var(--dim); } #status.ok{color:var(--ok)} #status.err{color:var(--err)} #status.warn{color:var(--accent)}
|
|
|
|
|
|
#readout { color:var(--dim); font-size:12px; }
|
|
|
|
|
|
|
|
|
|
|
|
/* the frame strip is the editing surface */
|
|
|
|
|
|
#strip { display:flex; flex-wrap:wrap; gap:4px; }
|
|
|
|
|
|
.fr { position:relative; cursor:pointer; border:2px solid transparent; border-radius:3px; line-height:0; }
|
|
|
|
|
|
.fr canvas { border-radius:1px; }
|
|
|
|
|
|
.fr span { position:absolute; left:2px; bottom:2px; font-size:10px; line-height:1.2;
|
|
|
|
|
|
padding:0 3px; border-radius:2px; background:#000a; color:#fff; }
|
|
|
|
|
|
.fr.keep { border-color:var(--ok); }
|
|
|
|
|
|
.fr.drop { border-color:#2a2f3e; }
|
|
|
|
|
|
.fr.drop canvas { opacity:.26; filter:grayscale(1); }
|
|
|
|
|
|
.fr.cur { border-color:var(--accent); }
|
2026-09-24 19:25:15 -04:00
|
|
|
|
.fr.cel::before { content:'▣'; position:absolute; left:3px; top:1px; color:#c084fc;
|
|
|
|
|
|
font-size:10px; text-shadow:0 0 2px #000; }
|
2026-09-24 14:51:15 -04:00
|
|
|
|
.fr.keep::after { content:'●'; position:absolute; right:3px; top:1px; color:var(--ok); font-size:10px; }
|
|
|
|
|
|
|
|
|
|
|
|
#sheet { display:flex; flex-wrap:wrap; gap:10px; }
|
|
|
|
|
|
#sheet .cell { display:flex; flex-direction:column; gap:3px; cursor:pointer; }
|
|
|
|
|
|
#sheet .cell span { color:var(--dim); font-size:11px; }
|
|
|
|
|
|
#sheet .cell:hover span { color:var(--accent); }
|
|
|
|
|
|
#palette { display:flex; gap:12px; flex-wrap:wrap; }
|
|
|
|
|
|
.sw { display:flex; align-items:center; gap:5px; color:var(--dim); font-size:11px; }
|
|
|
|
|
|
.sw input { width:26px; height:20px; padding:0; border:1px solid var(--line); background:none; }
|
|
|
|
|
|
.legend { color:var(--dim); font-size:11px; margin-top:6px; }
|
Paint: vector background cels on the kept frames
A sketch, and labelled as one. It exists to test whether the aesthetic holds
when a human draws the background rather than the tracker deriving it, and it
is meant to be replaced by a real paint surface with onion skin and undo. Kept
to one dependency-free module so throwing it away is a delete, not surgery.
Cels are drawn on the frames that get their own drawing and hold until the
next one - the same rule the plate follows, and literally the same lookup, so
the two cannot disagree about what is on screen. Editing always targets the
cel you can see, so you can scrub anywhere and keep drawing.
Pen places vertices and closes on the first one. Edit drags vertices or whole
shapes, shift-click inserts, alt-click removes. Layers stack front-at-top with
per-layer colour, show/hide and reorder. Drag a strip thumbnail onto the canvas
to seed this cel from that one, every layer, as a deep copy - sharing the point
arrays would make two cels silently edit each other.
Two rules are enforced rather than left to discipline:
Colours are PALETTE INDICES, never RGB. Sampling colour from the source is the
one move docs/design.md calls irrecoverable, and a paint tool is exactly where
that discipline would leak, so the picker cannot express a colour outside the
ramp.
Vertices snap to the 320x200 grid. On a hard-edged indexed rasteriser a shape
nudged by 0.4px moves an edge by a whole pixel or not at all depending on where
it happens to land, so sub-pixel vertices shimmer instead of holding still.
Drawings autosave to localStorage per take name. They are the only thing here a
person made by hand; everything else regenerates. Not in the .take export yet.
Also adds serve.py, a no-store dev server. python3 -m http.server sends
Last-Modified and browsers cache ES modules on it hard enough that a reload
serves a stale app.js against a fresh index.html: the new knobs appear, nothing
wires them, no error fires, and it reads as "your feature does not work". That
cost real time this session.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 19:08:07 -04:00
|
|
|
|
/* paint: a sketch, see js/paint.js */
|
|
|
|
|
|
#cv-paint { border:1px solid var(--line); cursor:crosshair; touch-action:none; }
|
|
|
|
|
|
#cv-paint.drop { border-color:var(--ok); }
|
|
|
|
|
|
#paintlayers { flex:0 0 280px; max-height:600px; overflow:auto; }
|
|
|
|
|
|
.lay { display:flex; align-items:center; gap:5px; padding:3px 4px; border-radius:3px;
|
|
|
|
|
|
border:1px solid transparent; cursor:pointer; }
|
|
|
|
|
|
.lay.sel { border-color:var(--accent); background:#1c2130; }
|
|
|
|
|
|
.lay i { width:12px; height:12px; border-radius:2px; border:1px solid #0006; flex:0 0 auto; }
|
|
|
|
|
|
.lay b { flex:1; font-weight:400; color:var(--dim); font-size:11px; }
|
|
|
|
|
|
.lay.sel b { color:var(--fg); }
|
|
|
|
|
|
.lay select { font:inherit; font-size:10px; background:#1c2130; color:var(--dim);
|
|
|
|
|
|
border:1px solid var(--line); border-radius:2px; padding:1px 2px; max-width:88px; }
|
|
|
|
|
|
.lay button { padding:0 5px; font-size:11px; line-height:18px; }
|
2026-09-24 14:51:15 -04:00
|
|
|
|
kbd { background:#1c2130; border:1px solid var(--line); border-radius:2px;
|
|
|
|
|
|
padding:0 4px; color:var(--fg); font-size:11px; }
|
2026-09-24 14:38:07 -04:00
|
|
|
|
</style>
|
|
|
|
|
|
</head>
|
|
|
|
|
|
<body>
|
|
|
|
|
|
<header>
|
Become arthur: a standalone suite, not an Animator Pro front-end
The test renderer turned out to be the product. Everything that decides how the
work looks - stabilisation, reduction, timing, frame removal, palette - already
happens here, and the flat indexed output already reads the way it should.
The reason to leave is in the original design's own rule: never make a timing
decision that requires a full render to evaluate. Honouring that moved every
judgement out of Animator Pro, which left the host doing nothing but writing a
file, in exchange for modal UI, minutes-long renders, one-level undo, FLX delta
invariants, a single tween state and a cel singleton.
What does NOT change is the constraint. 320x200, indexed palette, flat fills,
no antialiasing - inherited, but load-bearing rather than accidental. The
rasteriser writes palette indices and expands to RGBA only at the end precisely
so nothing can soften an edge. Modern conveniences belong in the workflow.
Adds docs/design.md: the principles, carried over without the Poco/FLX/cel
machinery, plus architecture and an honest list of what is missing - the
largest gap being that plates still have nowhere to be drawn.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:47:41 -04:00
|
|
|
|
<h1>arthur <span>— rotoscope + frame removal</span></h1>
|
2026-09-24 14:51:15 -04:00
|
|
|
|
<button id="btn-synth">Synthetic</button>
|
|
|
|
|
|
<input type="text" id="framedir" value="frames" size="7" title="frame directory">
|
2026-09-24 14:38:07 -04:00
|
|
|
|
<button id="btn-frames">Load frames</button>
|
|
|
|
|
|
<button id="btn-play">Play</button>
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
|
<select id="irisAnchor" title="what the iris hangs off">
|
|
|
|
|
|
<option value="steady" selected>iris: steady</option>
|
|
|
|
|
|
<option value="free">iris: free</option>
|
|
|
|
|
|
<option value="locked">iris: locked</option>
|
|
|
|
|
|
</select>
|
|
|
|
|
|
<select id="gazeOrigin" title="what counts as looking straight ahead">
|
|
|
|
|
|
<option value="median" selected>origin: median</option>
|
|
|
|
|
|
<option value="neutral">origin: neutral f</option>
|
|
|
|
|
|
</select>
|
|
|
|
|
|
<select id="exposure" title="exposure — how often the picture gets a new drawing">
|
|
|
|
|
|
<option value="1" selected>on 1s</option>
|
|
|
|
|
|
<option value="2">on 2s</option>
|
|
|
|
|
|
<option value="3">on 3s</option>
|
|
|
|
|
|
<option value="4">on 4s</option>
|
|
|
|
|
|
</select>
|
2026-09-24 14:53:09 -04:00
|
|
|
|
<select id="speed" title="playback speed"><option value="1">1x</option><option value="0.5">½x</option><option value="0.25">¼x</option></select>
|
|
|
|
|
|
<audio id="audio" controls hidden style="height:28px;vertical-align:middle"></audio>
|
2026-09-24 14:51:15 -04:00
|
|
|
|
<button id="btn-keepall">Keep all</button>
|
|
|
|
|
|
<button id="btn-suggest">Suggest</button>
|
|
|
|
|
|
<input type="text" id="takename" value="line_01" size="9" title="take name">
|
Registered photo underlay as the plate reference
The generated face oval was never going to be good enough to draw from:
MediaPipe's face oval is the FACE boundary, cut at the hairline and excluding
hair, ears, jaw and neck, so it is an egg by construction. Segmentation would
give a real head outline but costs a 16MB model and per-frame inference for a
shape that gets replaced by a drawing anyway.
So the plate layer becomes switchable, and the useful modes are photographic:
the source frame mapped into raster space through the same transform chain the
contours go through. Registration is the whole point - the head sits still and
a drawing traced from the underlay is already aligned to the mouth. An
unregistered underlay would be decoration.
- underlay.js: pixel->raster affine (a general affine, since MediaPipe
normalises x by width and y by height), registered draw, palette posterise
- plate modes: photo / photo dim / posterized / oval / oval+photo / none, B cycles
- worksheet cells are registered composites rather than raw crops
- Save frame 4x writes a 1280x800 PNG to draw on
- selftest: FACE_OVAL simplicity, which was never asserted; a wrong ordering
there reads as a lumpy plate rather than an obvious bowtie
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 14:57:59 -04:00
|
|
|
|
<select id="plateMode" title="plate representation (B to cycle)">
|
|
|
|
|
|
<option value="photo-dim">photo dim</option>
|
|
|
|
|
|
<option value="photo">photo</option>
|
|
|
|
|
|
<option value="posterize">posterized</option>
|
|
|
|
|
|
<option value="oval" selected>oval</option>
|
|
|
|
|
|
<option value="oval+photo">oval + photo</option>
|
|
|
|
|
|
<option value="off">none</option>
|
|
|
|
|
|
</select>
|
|
|
|
|
|
<button id="btn-saveframe">Save frame 4x</button>
|
2026-09-24 14:38:07 -04:00
|
|
|
|
<button id="btn-export">Export .take</button>
|
|
|
|
|
|
<span id="status"></span>
|
|
|
|
|
|
</header>
|
|
|
|
|
|
|
|
|
|
|
|
<main>
|
|
|
|
|
|
<div class="row">
|
|
|
|
|
|
<div class="panel">
|
|
|
|
|
|
<h2>source + landmarks</h2>
|
|
|
|
|
|
<canvas id="cv-source"></canvas>
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
|
<div class="legend">outer lip <b style="color:#4ade80">—</b> · inner lip <b style="color:#f87171">—</b> ·
|
Brows: traced ring, quantised raise
A brow at 320x200 is fourteen pixels wide and three tall. Its shape carries
almost nothing at that size; its height above the eye carries the expression,
and a brow raise is the most legible beat on a face. So the ring is traced and
the height is quantised - the split the eyes already got, where the lid is a
traced feature and the iris a quantised primitive.
The decomposition is the point. The traced ring already contains the real
height, so adding a quantised raise on top would move the brow twice. The
height is measured OUT of the ring, quantised, and put back, so the shape that
renders is his at a height that snaps between a few levels and holds.
Measured at both ends rather than as one number, because raise and tilt are
different expressions out of one mechanism: both ends up is surprise, inner up
alone is worry, inner down is anger. They share a dwell - the gaze quantiser,
renamed quantizeSnap now that it has two callers - so the brow hits its pose in
one frame instead of crawling into it with one end arriving before the other.
Measured against the eye's corner midpoint, never its lid. Same trap the gaze
origin has and worth avoiding twice: brows and lids move together constantly,
so a brow that jumped on every blink would read as a tic. Rest pose from the
take median rather than the neutral frame, for the reason gaze learned the hard
way - that frame is picked by minimum mouth aperture and says nothing about the
brows.
Two correspondences resolved from geometry, not declared: which ring is which
brow, and which end is the outer one. The second matters more - backwards, the
tilt mirrors and worry renders as its own opposite, which reads as a directed
performance choice rather than a bug and would never be questioned. Which EDGE
is upper is deliberately left unresolved: it traverses the same ring the other
way, an even-odd fill has no winding, and both ends still land on fixed slots.
Also fixes a bug from the exposure work: the live render applied exposure to
the plate and the mouth but not to the eyes, so on 2s the preview and the
export disagreed. A preview that disagrees with the export is the one bug this
tool cannot afford. perfIndex now exists as a named thing so the two paths
cannot drift apart again.
91 -> 105 assertions. Ground truth on all four synthetic brow poses, tilt
separating worry from anger by sign, a blink not faking a raise, and a shared
dwell never emitting a half-raised brow.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:11:25 -04:00
|
|
|
|
lids <b style="color:#60a5fa">—</b> · iris <b style="color:#fbbf24">—</b> ·
|
|
|
|
|
|
brows <b style="color:#c084fc">—</b></div>
|
2026-09-24 14:38:07 -04:00
|
|
|
|
</div>
|
|
|
|
|
|
<div class="panel">
|
|
|
|
|
|
<h2>stabilised (head-local)</h2>
|
|
|
|
|
|
<canvas id="cv-stab"></canvas>
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
|
<div class="legend">should sit still except the mouth and eyes · grey = held plate outline<br>
|
|
|
|
|
|
dark green ghost = unshifted mouth when lead ≠ 0 · red lid ring = blink</div>
|
2026-09-24 14:38:07 -04:00
|
|
|
|
</div>
|
|
|
|
|
|
<div class="panel">
|
|
|
|
|
|
<h2>flat render — 320×200 indexed</h2>
|
|
|
|
|
|
<canvas id="cv-render"></canvas>
|
|
|
|
|
|
<div class="legend" id="framelabel"></div>
|
Registered photo underlay as the plate reference
The generated face oval was never going to be good enough to draw from:
MediaPipe's face oval is the FACE boundary, cut at the hairline and excluding
hair, ears, jaw and neck, so it is an egg by construction. Segmentation would
give a real head outline but costs a 16MB model and per-frame inference for a
shape that gets replaced by a drawing anyway.
So the plate layer becomes switchable, and the useful modes are photographic:
the source frame mapped into raster space through the same transform chain the
contours go through. Registration is the whole point - the head sits still and
a drawing traced from the underlay is already aligned to the mouth. An
unregistered underlay would be decoration.
- underlay.js: pixel->raster affine (a general affine, since MediaPipe
normalises x by width and y by height), registered draw, palette posterise
- plate modes: photo / photo dim / posterized / oval / oval+photo / none, B cycles
- worksheet cells are registered composites rather than raw crops
- Save frame 4x writes a 1280x800 PNG to draw on
- selftest: FACE_OVAL simplicity, which was never asserted; a wrong ordering
there reads as a lumpy plate rather than an obvious bowtie
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 14:57:59 -04:00
|
|
|
|
<div class="legend">audio drives the clock — dropped frames, never drift<br>
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
|
<b>exposure</b> holds the picture on a grid: rip at 24, render on 2s for
|
|
|
|
|
|
12. The dense track and the audio are untouched, so it is reversible<br>
|
Registered photo underlay as the plate reference
The generated face oval was never going to be good enough to draw from:
MediaPipe's face oval is the FACE boundary, cut at the hairline and excluding
hair, ears, jaw and neck, so it is an egg by construction. Segmentation would
give a real head outline but costs a 16MB model and per-frame inference for a
shape that gets replaced by a drawing anyway.
So the plate layer becomes switchable, and the useful modes are photographic:
the source frame mapped into raster space through the same transform chain the
contours go through. Registration is the whole point - the head sits still and
a drawing traced from the underlay is already aligned to the mouth. An
unregistered underlay would be decoration.
- underlay.js: pixel->raster affine (a general affine, since MediaPipe
normalises x by width and y by height), registered draw, palette posterise
- plate modes: photo / photo dim / posterized / oval / oval+photo / none, B cycles
- worksheet cells are registered composites rather than raw crops
- Save frame 4x writes a 1280x800 PNG to draw on
- selftest: FACE_OVAL simplicity, which was never asserted; a wrong ordering
there reads as a lumpy plate rather than an obvious bowtie
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 14:57:59 -04:00
|
|
|
|
plate representation: <kbd>B</kbd> cycles · photo modes are <b>registered</b>
|
|
|
|
|
|
into raster space, so tracing them lands on the mouth</div>
|
2026-09-24 14:38:07 -04:00
|
|
|
|
</div>
|
|
|
|
|
|
</div>
|
|
|
|
|
|
|
2026-09-24 14:51:15 -04:00
|
|
|
|
<div class="panel">
|
|
|
|
|
|
<h2>frames — green kept (gets its own drawing) · dim held from the last kept frame</h2>
|
|
|
|
|
|
<div id="strip"></div>
|
|
|
|
|
|
<input type="range" id="scrub" min="0" max="0" value="0" style="width:100%;margin-top:10px">
|
|
|
|
|
|
<div class="legend">
|
Registered photo underlay as the plate reference
The generated face oval was never going to be good enough to draw from:
MediaPipe's face oval is the FACE boundary, cut at the hairline and excluding
hair, ears, jaw and neck, so it is an egg by construction. Segmentation would
give a real head outline but costs a 16MB model and per-frame inference for a
shape that gets replaced by a drawing anyway.
So the plate layer becomes switchable, and the useful modes are photographic:
the source frame mapped into raster space through the same transform chain the
contours go through. Registration is the whole point - the head sits still and
a drawing traced from the underlay is already aligned to the mouth. An
unregistered underlay would be decoration.
- underlay.js: pixel->raster affine (a general affine, since MediaPipe
normalises x by width and y by height), registered draw, palette posterise
- plate modes: photo / photo dim / posterized / oval / oval+photo / none, B cycles
- worksheet cells are registered composites rather than raw crops
- Save frame 4x writes a 1280x800 PNG to draw on
- selftest: FACE_OVAL simplicity, which was never asserted; a wrong ordering
there reads as a lumpy plate rather than an obvious bowtie
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 14:57:59 -04:00
|
|
|
|
<kbd>←</kbd> <kbd>→</kbd> step · <kbd>X</kbd> delete · <kbd>K</kbd> keep · <kbd>B</kbd> background ·
|
2026-09-24 15:34:08 -04:00
|
|
|
|
<kbd>[</kbd> <kbd>]</kbd> lead ·
|
2026-09-24 14:51:15 -04:00
|
|
|
|
click to select · double-click or shift-click to toggle ·
|
|
|
|
|
|
the mouth keeps <b>every</b> frame regardless
|
|
|
|
|
|
</div>
|
|
|
|
|
|
<div class="legend" id="readout"></div>
|
|
|
|
|
|
</div>
|
|
|
|
|
|
|
2026-09-24 14:38:07 -04:00
|
|
|
|
<div class="row">
|
2026-09-24 14:51:15 -04:00
|
|
|
|
<div class="panel" style="flex:1 1 400px">
|
2026-09-24 14:38:07 -04:00
|
|
|
|
<h2>knobs</h2>
|
|
|
|
|
|
<label class="ctl"><span>vertices</span><input type="range" id="verts" min="4" max="16" step="2" value="8"><output id="vertsv"></output></label>
|
2026-09-24 15:34:08 -04:00
|
|
|
|
<label class="ctl"><span>mouth lead ±f</span><input type="range" id="lead" min="-6" max="6" value="0"><output id="leadv"></output></label>
|
2026-09-24 14:51:15 -04:00
|
|
|
|
<label class="ctl"><span>contour avg ±f</span><input type="range" id="contourSmooth" min="0" max="4" value="1"><output id="contourSmoothv"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>anchor avg ±f</span><input type="range" id="smoothWin" min="0" max="8" value="2"><output id="smoothWinv"></output></label>
|
2026-09-24 14:38:07 -04:00
|
|
|
|
<label class="ctl"><span>closed-mouth cut</span><input type="range" id="apertureThresh" min="0" max="400" value="120"><output id="apertureThreshv"></output></label>
|
Fix teeth band filling the whole mouth
Three causes, all of them mine:
Otsu always returns a split, including on a homogeneous region - given a dark
cavity with no teeth it invents a threshold and calls half the pixels bright.
The gate is now the separation between the two class means, which is the only
thing that says whether the split means anything. Coverage was the wrong
signal: it is high both when the mouth is full of teeth and when the region is
uniformly dark and Otsu has split noise.
The row scan tracked the last qualifying row anywhere rather than where the run
from the top stops, so one bright row near the bottom - a lit lower lip inside
the ring - pushed the line to full height. It now breaks at the first failing
row once the run has started.
MediaPipe's inner lip landmarks sit slightly outside the real opening, so the
sampled region included lip pixels, which are bright and sit exactly at the
boundary where they do most damage. The ring is now eroded toward its centroid
before sampling, with the amount exposed as a knob.
Adds a diagnostic panel showing the sampled crop, pixels above threshold, and
the resolved line, because tuning this from numbers alone does not tell you
whether the region being measured is even the right region.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:14:49 -04:00
|
|
|
|
<label class="ctl"><span>teeth contrast</span><input type="range" id="teethOn" min="1" max="60" value="16"><output id="teethOnv"></output></label>
|
Teeth as an extracted blob contour, not a clipped band
The band filled the mouth because a band is the wrong reduction: the bright
region is a blob, and reading it as "everything above a line" throws the shape
away.
Extracting a contour reintroduces the vertex-correspondence problem that made
me avoid it, but for a blob there is a way out. Radial sampling from the
centroid along N fixed directions makes vertex k always mean "the extent in
direction k": correspondence holds by construction, the count is fixed, and
temporal smoothing cannot reorder anything. It also yields a star-shaped
reduction, which suits flat colour.
Tongue rejection, which the band had no way to express:
- pixels red relative to their own brightness are dropped (teeth are neutral)
- component choice is biased toward the top of the cavity, since area alone
picks the tongue when the mouth is wide
- separate inner and outer controls: cavity erode pulls the sampled region off
the lip edge, blob grow/erode resizes the found blob
Also: a knob wired in app.js but missing from index.html threw during wiring
and left a blank page with nothing useful in the console - which is exactly
what happened to teethDwell in the previous commit. el() now names the missing
id, and window.onerror surfaces it in the status line.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:24:52 -04:00
|
|
|
|
<label class="ctl"><span>cavity erode</span><input type="range" id="teethErode" min="0" max="45" value="18"><output id="teethErodev"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>blob grow/erode</span><input type="range" id="blobGrow" min="-4" max="4" value="0"><output id="blobGrowv"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>tongue reject</span><input type="range" id="tongueReject" min="2" max="40" value="18"><output id="tongueRejectv"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>prefer upper</span><input type="range" id="topBias" min="0" max="120" value="60"><output id="topBiasv"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>teeth vertices</span><input type="range" id="teethVerts" min="5" max="20" value="10"><output id="teethVertsv"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>teeth avg ±f</span><input type="range" id="teethSmooth" min="0" max="4" value="1"><output id="teethSmoothv"></output></label>
|
2026-09-24 15:12:09 -04:00
|
|
|
|
<label class="ctl"><span>teeth dwell</span><input type="range" id="teethDwell" min="0" max="6" value="1"><output id="teethDwellv"></output></label>
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
|
<label class="ctl"><span>eye vertices</span><input type="range" id="eyeVerts" min="4" max="12" step="2" value="8"><output id="eyeVertsv"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>lash line</span><input type="range" id="lashPx" min="0" max="3" value="1"><output id="lashPxv"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>blink cut</span><input type="range" id="blinkCut" min="20" max="300" value="130"><output id="blinkCutv"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>blink hold ±f</span><input type="range" id="blinkHold" min="1" max="5" value="2"><output id="blinkHoldv"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>blink dwell</span><input type="range" id="blinkDwell" min="0" max="4" value="0"><output id="blinkDwellv"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>gaze gain</span><input type="range" id="gazeGain" min="50" max="400" value="100"><output id="gazeGainv"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>gaze step</span><input type="range" id="gazeStep" min="0" max="6" value="2"><output id="gazeStepv"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>gaze dwell</span><input type="range" id="gazeDwell" min="0" max="6" value="2"><output id="gazeDwellv"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>iris size</span><input type="range" id="irisSize" min="20" max="70" value="42"><output id="irisSizev"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>pupil</span><input type="range" id="pupilPx" min="0" max="7" value="3"><output id="pupilPxv"></output></label>
|
Brows: traced ring, quantised raise
A brow at 320x200 is fourteen pixels wide and three tall. Its shape carries
almost nothing at that size; its height above the eye carries the expression,
and a brow raise is the most legible beat on a face. So the ring is traced and
the height is quantised - the split the eyes already got, where the lid is a
traced feature and the iris a quantised primitive.
The decomposition is the point. The traced ring already contains the real
height, so adding a quantised raise on top would move the brow twice. The
height is measured OUT of the ring, quantised, and put back, so the shape that
renders is his at a height that snaps between a few levels and holds.
Measured at both ends rather than as one number, because raise and tilt are
different expressions out of one mechanism: both ends up is surprise, inner up
alone is worry, inner down is anger. They share a dwell - the gaze quantiser,
renamed quantizeSnap now that it has two callers - so the brow hits its pose in
one frame instead of crawling into it with one end arriving before the other.
Measured against the eye's corner midpoint, never its lid. Same trap the gaze
origin has and worth avoiding twice: brows and lids move together constantly,
so a brow that jumped on every blink would read as a tic. Rest pose from the
take median rather than the neutral frame, for the reason gaze learned the hard
way - that frame is picked by minimum mouth aperture and says nothing about the
brows.
Two correspondences resolved from geometry, not declared: which ring is which
brow, and which end is the outer one. The second matters more - backwards, the
tilt mirrors and worry renders as its own opposite, which reads as a directed
performance choice rather than a bug and would never be questioned. Which EDGE
is upper is deliberately left unresolved: it traverses the same ring the other
way, an even-odd fill has no winding, and both ends still land on fixed slots.
Also fixes a bug from the exposure work: the live render applied exposure to
the plate and the mouth but not to the eyes, so on 2s the preview and the
export disagreed. A preview that disagrees with the export is the one bug this
tool cannot afford. perfIndex now exists as a named thing so the two paths
cannot drift apart again.
91 -> 105 assertions. Ground truth on all four synthetic brow poses, tilt
separating worry from anger by sign, a blink not faking a raise, and a shared
dwell never emitting a half-raised brow.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:11:25 -04:00
|
|
|
|
<label class="ctl"><span>brow vertices</span><input type="range" id="browVerts" min="4" max="10" step="2" value="6"><output id="browVertsv"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>brow weight</span><input type="range" id="browWeight" min="0" max="3" value="1"><output id="browWeightv"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>brow raise gain</span><input type="range" id="browGain" min="50" max="300" value="100"><output id="browGainv"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>brow step</span><input type="range" id="browStep" min="0" max="6" value="2"><output id="browStepv"></output></label>
|
|
|
|
|
|
<label class="ctl"><span>brow dwell</span><input type="range" id="browDwell" min="0" max="6" value="2"><output id="browDwellv"></output></label>
|
2026-09-24 14:51:15 -04:00
|
|
|
|
<label class="ctl"><span>suggest tolerance</span><input type="range" id="tol" min="2" max="60" value="14"><output id="tolv"></output></label>
|
|
|
|
|
|
<div class="legend">
|
2026-09-24 15:34:08 -04:00
|
|
|
|
<b>mouth lead</b> shifts the performance tracks earlier (positive) against
|
|
|
|
|
|
the audio and the head. Averaging has no phase lag but it blurs onsets, so
|
|
|
|
|
|
an opening reads later than it is; animators also draw mouths a frame or
|
|
|
|
|
|
two ahead of the sound as standard practice. <kbd>[</kbd> <kbd>]</kbd>.<br>
|
2026-09-24 14:51:15 -04:00
|
|
|
|
<b>contour avg</b> 0 = off, 1 = ±1 frame. Removes per-frame landmark
|
|
|
|
|
|
jitter. Push past 2 and it starts eating articulation.<br>
|
|
|
|
|
|
<b>anchor avg</b> smooths the head transform only — never the contour.<br>
|
Teeth as an extracted blob contour, not a clipped band
The band filled the mouth because a band is the wrong reduction: the bright
region is a blob, and reading it as "everything above a line" throws the shape
away.
Extracting a contour reintroduces the vertex-correspondence problem that made
me avoid it, but for a blob there is a way out. Radial sampling from the
centroid along N fixed directions makes vertex k always mean "the extent in
direction k": correspondence holds by construction, the count is fixed, and
temporal smoothing cannot reorder anything. It also yields a star-shaped
reduction, which suits flat colour.
Tongue rejection, which the band had no way to express:
- pixels red relative to their own brightness are dropped (teeth are neutral)
- component choice is biased toward the top of the cavity, since area alone
picks the tongue when the mouth is wide
- separate inner and outer controls: cavity erode pulls the sampled region off
the lip edge, blob grow/erode resizes the found blob
Also: a knob wired in app.js but missing from index.html threw during wiring
and left a blank page with nothing useful in the console - which is exactly
what happened to teethDwell in the previous commit. el() now names the missing
id, and window.onerror surfaces it in the status line.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:24:52 -04:00
|
|
|
|
<b>teeth contrast</b> gates on how far apart the cavity's dark and bright
|
|
|
|
|
|
halves are — Otsu always returns <i>some</i> threshold, so this is what
|
|
|
|
|
|
stops it inventing teeth in a dark mouth.
|
|
|
|
|
|
<b>cavity erode</b> pulls the sampled region in from the lip edge;
|
|
|
|
|
|
<b>blob grow/erode</b> resizes the found blob itself.
|
|
|
|
|
|
<b>tongue reject</b> drops pixels that are red relative to their own
|
|
|
|
|
|
brightness; <b>prefer upper</b> biases component choice toward the top of
|
|
|
|
|
|
the cavity, where teeth are and the tongue is not.
|
|
|
|
|
|
<b>dwell</b> is how many frames a presence change must persist.<br>
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
|
<b>blink cut</b> is lid gap over corner distance — normalised, so one
|
|
|
|
|
|
value carries across takes. <b>blink hold</b> is the minimum length of a
|
|
|
|
|
|
blink: a real blink is one frame at 12fps and a single frame of closed
|
|
|
|
|
|
eye reads as a dropout, so it is extended to a beat.
|
Brows: traced ring, quantised raise
A brow at 320x200 is fourteen pixels wide and three tall. Its shape carries
almost nothing at that size; its height above the eye carries the expression,
and a brow raise is the most legible beat on a face. So the ring is traced and
the height is quantised - the split the eyes already got, where the lid is a
traced feature and the iris a quantised primitive.
The decomposition is the point. The traced ring already contains the real
height, so adding a quantised raise on top would move the brow twice. The
height is measured OUT of the ring, quantised, and put back, so the shape that
renders is his at a height that snaps between a few levels and holds.
Measured at both ends rather than as one number, because raise and tilt are
different expressions out of one mechanism: both ends up is surprise, inner up
alone is worry, inner down is anger. They share a dwell - the gaze quantiser,
renamed quantizeSnap now that it has two callers - so the brow hits its pose in
one frame instead of crawling into it with one end arriving before the other.
Measured against the eye's corner midpoint, never its lid. Same trap the gaze
origin has and worth avoiding twice: brows and lids move together constantly,
so a brow that jumped on every blink would read as a tic. Rest pose from the
take median rather than the neutral frame, for the reason gaze learned the hard
way - that frame is picked by minimum mouth aperture and says nothing about the
brows.
Two correspondences resolved from geometry, not declared: which ring is which
brow, and which end is the outer one. The second matters more - backwards, the
tilt mirrors and worry renders as its own opposite, which reads as a directed
performance choice rather than a bug and would never be questioned. Which EDGE
is upper is deliberately left unresolved: it traverses the same ring the other
way, an even-odd fill has no winding, and both ends still land on fixed slots.
Also fixes a bug from the exposure work: the live render applied exposure to
the plate and the mouth but not to the eyes, so on 2s the preview and the
export disagreed. A preview that disagrees with the export is the one bug this
tool cannot afford. perfIndex now exists as a named thing so the two paths
cannot drift apart again.
91 -> 105 assertions. Ground truth on all four synthetic brow poses, tilt
separating worry from anger by sign, a blink not faking a raise, and a shared
dwell never emitting a half-raised brow.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:11:25 -04:00
|
|
|
|
<b>brow step</b> and <b>brow dwell</b> quantise the brow's HEIGHT above
|
|
|
|
|
|
the eye, not its shape — the traced ring is his, the height snaps between
|
|
|
|
|
|
a few levels and holds. Both ends move independently, so raise and tilt
|
|
|
|
|
|
come out of one control: both up is surprise, inner up is worry, inner
|
|
|
|
|
|
down is anger. <b>brow weight</b> thickens the ring, which it needs at
|
|
|
|
|
|
three pixels tall.<br>
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
|
<b>pupil</b> is a square, in whole pixels, 0 to turn it off: at this size
|
|
|
|
|
|
a circle of radius 1.5 is a plus sign with the corners gnawed off and it
|
|
|
|
|
|
changes shape as it moves, where a square stays the mark you drew.<br>
|
|
|
|
|
|
<b>gaze step</b> is the grid the iris snaps to, in raster pixels, and
|
|
|
|
|
|
<b>gaze dwell</b> is how long a new cell must hold — together they turn
|
|
|
|
|
|
drift into saccades. <b>gaze gain</b> exaggerates or damps the throw;
|
|
|
|
|
|
measured excursion is small and a character usually wants more of it.<br>
|
2026-09-24 14:51:15 -04:00
|
|
|
|
<b>suggest tolerance</b> only affects the Suggest button: max head movement
|
|
|
|
|
|
allowed before a new drawing is required.
|
|
|
|
|
|
</div>
|
2026-09-24 14:38:07 -04:00
|
|
|
|
</div>
|
Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.
Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.
Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.
Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.
The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.
Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.
The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.
Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.
Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.
Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.
41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
|
|
|
|
<div class="panel" style="flex:0 1 170px">
|
|
|
|
|
|
<h2>gaze field</h2>
|
|
|
|
|
|
<canvas id="cv-gaze"></canvas>
|
|
|
|
|
|
<div class="legend" id="eyeinfo" style="white-space:pre-line"></div>
|
|
|
|
|
|
<div class="legend">green = every cell the iris visits in the take ·
|
|
|
|
|
|
grey = raw · amber = where it is now, quantised</div>
|
|
|
|
|
|
</div>
|
Fix teeth band filling the whole mouth
Three causes, all of them mine:
Otsu always returns a split, including on a homogeneous region - given a dark
cavity with no teeth it invents a threshold and calls half the pixels bright.
The gate is now the separation between the two class means, which is the only
thing that says whether the split means anything. Coverage was the wrong
signal: it is high both when the mouth is full of teeth and when the region is
uniformly dark and Otsu has split noise.
The row scan tracked the last qualifying row anywhere rather than where the run
from the top stops, so one bright row near the bottom - a lit lower lip inside
the ring - pushed the line to full height. It now breaks at the first failing
row once the run has started.
MediaPipe's inner lip landmarks sit slightly outside the real opening, so the
sampled region included lip pixels, which are bright and sit exactly at the
boundary where they do most damage. The ring is now eroded toward its centroid
before sampling, with the amount exposed as a knob.
Adds a diagnostic panel showing the sampled crop, pixels above threshold, and
the resolved line, because tuning this from numbers alone does not tell you
whether the region being measured is even the right region.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:14:49 -04:00
|
|
|
|
<div class="panel" style="flex:0 1 190px">
|
|
|
|
|
|
<h2>teeth measurement</h2>
|
|
|
|
|
|
<div id="cv-teeth"></div>
|
|
|
|
|
|
<div class="legend" id="teethinfo"></div>
|
Teeth as an extracted blob contour, not a clipped band
The band filled the mouth because a band is the wrong reduction: the bright
region is a blob, and reading it as "everything above a line" throws the shape
away.
Extracting a contour reintroduces the vertex-correspondence problem that made
me avoid it, but for a blob there is a way out. Radial sampling from the
centroid along N fixed directions makes vertex k always mean "the extent in
direction k": correspondence holds by construction, the count is fixed, and
temporal smoothing cannot reorder anything. It also yields a star-shaped
reduction, which suits flat colour.
Tongue rejection, which the band had no way to express:
- pixels red relative to their own brightness are dropped (teeth are neutral)
- component choice is biased toward the top of the cavity, since area alone
picks the tongue when the mouth is wide
- separate inner and outer controls: cavity erode pulls the sampled region off
the lip edge, blob grow/erode resizes the found blob
Also: a knob wired in app.js but missing from index.html threw during wiring
and left a blank page with nothing useful in the console - which is exactly
what happened to teethDwell in the previous commit. el() now names the missing
id, and window.onerror surfaces it in the status line.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:24:52 -04:00
|
|
|
|
<div class="legend">green = kept pixels · amber = extracted contour</div>
|
Fix teeth band filling the whole mouth
Three causes, all of them mine:
Otsu always returns a split, including on a homogeneous region - given a dark
cavity with no teeth it invents a threshold and calls half the pixels bright.
The gate is now the separation between the two class means, which is the only
thing that says whether the split means anything. Coverage was the wrong
signal: it is high both when the mouth is full of teeth and when the region is
uniformly dark and Otsu has split noise.
The row scan tracked the last qualifying row anywhere rather than where the run
from the top stops, so one bright row near the bottom - a lit lower lip inside
the ring - pushed the line to full height. It now breaks at the first failing
row once the run has started.
MediaPipe's inner lip landmarks sit slightly outside the real opening, so the
sampled region included lip pixels, which are bright and sit exactly at the
boundary where they do most damage. The ring is now eroded toward its centroid
before sampling, with the amount exposed as a knob.
Adds a diagnostic panel showing the sampled crop, pixels above threshold, and
the resolved line, because tuning this from numbers alone does not tell you
whether the region being measured is even the right region.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 15:14:49 -04:00
|
|
|
|
</div>
|
|
|
|
|
|
<div class="panel" style="flex:1 1 240px">
|
2026-09-24 14:38:07 -04:00
|
|
|
|
<h2>palette</h2>
|
|
|
|
|
|
<div id="palette"></div>
|
|
|
|
|
|
<div class="legend" style="margin-top:12px">
|
|
|
|
|
|
Flat indexed fills, no antialiasing — the rasteriser writes palette
|
|
|
|
|
|
indices, the way <code>csd_render_poly</code> does.
|
|
|
|
|
|
</div>
|
|
|
|
|
|
</div>
|
|
|
|
|
|
</div>
|
|
|
|
|
|
|
Paint: vector background cels on the kept frames
A sketch, and labelled as one. It exists to test whether the aesthetic holds
when a human draws the background rather than the tracker deriving it, and it
is meant to be replaced by a real paint surface with onion skin and undo. Kept
to one dependency-free module so throwing it away is a delete, not surgery.
Cels are drawn on the frames that get their own drawing and hold until the
next one - the same rule the plate follows, and literally the same lookup, so
the two cannot disagree about what is on screen. Editing always targets the
cel you can see, so you can scrub anywhere and keep drawing.
Pen places vertices and closes on the first one. Edit drags vertices or whole
shapes, shift-click inserts, alt-click removes. Layers stack front-at-top with
per-layer colour, show/hide and reorder. Drag a strip thumbnail onto the canvas
to seed this cel from that one, every layer, as a deep copy - sharing the point
arrays would make two cels silently edit each other.
Two rules are enforced rather than left to discipline:
Colours are PALETTE INDICES, never RGB. Sampling colour from the source is the
one move docs/design.md calls irrecoverable, and a paint tool is exactly where
that discipline would leak, so the picker cannot express a colour outside the
ramp.
Vertices snap to the 320x200 grid. On a hard-edged indexed rasteriser a shape
nudged by 0.4px moves an edge by a whole pixel or not at all depending on where
it happens to land, so sub-pixel vertices shimmer instead of holding still.
Drawings autosave to localStorage per take name. They are the only thing here a
person made by hand; everything else regenerates. Not in the .take export yet.
Also adds serve.py, a no-store dev server. python3 -m http.server sends
Last-Modified and browsers cache ES modules on it hard enough that a reload
serves a stale app.js against a fresh index.html: the new knobs appear, nothing
wires them, no error fires, and it reads as "your feature does not work". That
cost real time this session.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 19:08:07 -04:00
|
|
|
|
<div class="panel">
|
|
|
|
|
|
<h2>paint — background cels, drawn on the frames that get their own drawing</h2>
|
|
|
|
|
|
<div style="display:flex;gap:8px;align-items:center;flex-wrap:wrap;margin-bottom:8px">
|
|
|
|
|
|
<select id="paintTool" title="tool">
|
|
|
|
|
|
<option value="pen" selected>pen</option>
|
|
|
|
|
|
<option value="edit">edit</option>
|
|
|
|
|
|
</select>
|
|
|
|
|
|
<select id="paintColor" title="palette colour — indices only, never an RGB value"></select>
|
|
|
|
|
|
<button id="btn-celprev">Copy previous</button>
|
|
|
|
|
|
<button id="btn-celclear">Clear cel</button>
|
|
|
|
|
|
<span class="legend" id="paintframe" style="margin:0"></span>
|
|
|
|
|
|
</div>
|
|
|
|
|
|
<div style="display:flex;gap:12px;flex-wrap:wrap">
|
|
|
|
|
|
<canvas id="cv-paint"></canvas>
|
|
|
|
|
|
<div id="paintlayers"></div>
|
|
|
|
|
|
</div>
|
|
|
|
|
|
<div class="legend" id="paintinfo"></div>
|
2026-09-24 19:25:15 -04:00
|
|
|
|
<div class="legend" id="paintcels"></div>
|
Paint: vector background cels on the kept frames
A sketch, and labelled as one. It exists to test whether the aesthetic holds
when a human draws the background rather than the tracker deriving it, and it
is meant to be replaced by a real paint surface with onion skin and undo. Kept
to one dependency-free module so throwing it away is a delete, not surgery.
Cels are drawn on the frames that get their own drawing and hold until the
next one - the same rule the plate follows, and literally the same lookup, so
the two cannot disagree about what is on screen. Editing always targets the
cel you can see, so you can scrub anywhere and keep drawing.
Pen places vertices and closes on the first one. Edit drags vertices or whole
shapes, shift-click inserts, alt-click removes. Layers stack front-at-top with
per-layer colour, show/hide and reorder. Drag a strip thumbnail onto the canvas
to seed this cel from that one, every layer, as a deep copy - sharing the point
arrays would make two cels silently edit each other.
Two rules are enforced rather than left to discipline:
Colours are PALETTE INDICES, never RGB. Sampling colour from the source is the
one move docs/design.md calls irrecoverable, and a paint tool is exactly where
that discipline would leak, so the picker cannot express a colour outside the
ramp.
Vertices snap to the 320x200 grid. On a hard-edged indexed rasteriser a shape
nudged by 0.4px moves an edge by a whole pixel or not at all depending on where
it happens to land, so sub-pixel vertices shimmer instead of holding still.
Drawings autosave to localStorage per take name. They are the only thing here a
person made by hand; everything else regenerates. Not in the .take export yet.
Also adds serve.py, a no-store dev server. python3 -m http.server sends
Last-Modified and browsers cache ES modules on it hard enough that a reload
serves a stale app.js against a fresh index.html: the new knobs appear, nothing
wires them, no error fires, and it reads as "your feature does not work". That
cost real time this session.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 19:08:07 -04:00
|
|
|
|
<div class="legend">
|
|
|
|
|
|
<b>pen</b> click to place vertices · click the green box on the first one, or
|
|
|
|
|
|
<kbd>Enter</kbd> / double-click, to close · <kbd>Esc</kbd> cancels ·
|
|
|
|
|
|
<kbd>Backspace</kbd> drops the last point<br>
|
|
|
|
|
|
<b>edit</b> click a shape to select · drag a vertex or the shape itself ·
|
|
|
|
|
|
<kbd>Shift</kbd>-click an edge inserts a vertex · <kbd>Alt</kbd>-click a vertex
|
|
|
|
|
|
removes it · <kbd>Del</kbd> deletes the layer<br>
|
2026-09-24 19:25:15 -04:00
|
|
|
|
<b>Copy previous</b> takes the last enabled frame that actually has a drawing
|
|
|
|
|
|
— skipping empty ones, so it behaves the same before and after you thin the
|
|
|
|
|
|
strip out. Frames carrying a drawing are marked <b style="color:#c084fc">▣</b>.<br>
|
Paint: vector background cels on the kept frames
A sketch, and labelled as one. It exists to test whether the aesthetic holds
when a human draws the background rather than the tracker deriving it, and it
is meant to be replaced by a real paint surface with onion skin and undo. Kept
to one dependency-free module so throwing it away is a delete, not surgery.
Cels are drawn on the frames that get their own drawing and hold until the
next one - the same rule the plate follows, and literally the same lookup, so
the two cannot disagree about what is on screen. Editing always targets the
cel you can see, so you can scrub anywhere and keep drawing.
Pen places vertices and closes on the first one. Edit drags vertices or whole
shapes, shift-click inserts, alt-click removes. Layers stack front-at-top with
per-layer colour, show/hide and reorder. Drag a strip thumbnail onto the canvas
to seed this cel from that one, every layer, as a deep copy - sharing the point
arrays would make two cels silently edit each other.
Two rules are enforced rather than left to discipline:
Colours are PALETTE INDICES, never RGB. Sampling colour from the source is the
one move docs/design.md calls irrecoverable, and a paint tool is exactly where
that discipline would leak, so the picker cannot express a colour outside the
ramp.
Vertices snap to the 320x200 grid. On a hard-edged indexed rasteriser a shape
nudged by 0.4px moves an edge by a whole pixel or not at all depending on where
it happens to land, so sub-pixel vertices shimmer instead of holding still.
Drawings autosave to localStorage per take name. They are the only thing here a
person made by hand; everything else regenerates. Not in the .take export yet.
Also adds serve.py, a no-store dev server. python3 -m http.server sends
Last-Modified and browsers cache ES modules on it hard enough that a reload
serves a stale app.js against a fresh index.html: the new knobs appear, nothing
wires them, no error fires, and it reads as "your feature does not work". That
cost real time this session.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 19:08:07 -04:00
|
|
|
|
<b>drag a frame</b> from the strip onto the canvas to copy its drawing, every
|
|
|
|
|
|
layer. Vertices snap to the 320×200 grid, and colours are palette indices —
|
|
|
|
|
|
you cannot pick one that is not in the ramp.
|
|
|
|
|
|
</div>
|
|
|
|
|
|
</div>
|
|
|
|
|
|
|
2026-09-24 14:38:07 -04:00
|
|
|
|
<div class="panel">
|
Registered photo underlay as the plate reference
The generated face oval was never going to be good enough to draw from:
MediaPipe's face oval is the FACE boundary, cut at the hairline and excluding
hair, ears, jaw and neck, so it is an egg by construction. Segmentation would
give a real head outline but costs a 16MB model and per-frame inference for a
shape that gets replaced by a drawing anyway.
So the plate layer becomes switchable, and the useful modes are photographic:
the source frame mapped into raster space through the same transform chain the
contours go through. Registration is the whole point - the head sits still and
a drawing traced from the underlay is already aligned to the mouth. An
unregistered underlay would be decoration.
- underlay.js: pixel->raster affine (a general affine, since MediaPipe
normalises x by width and y by height), registered draw, palette posterise
- plate modes: photo / photo dim / posterized / oval / oval+photo / none, B cycles
- worksheet cells are registered composites rather than raw crops
- Save frame 4x writes a 1280x800 PNG to draw on
- selftest: FACE_OVAL simplicity, which was never asserted; a wrong ordering
there reads as a lumpy plate rather than an obvious bowtie
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 14:57:59 -04:00
|
|
|
|
<h2>drawings needed — registered reference per kept frame, with the range it holds</h2>
|
2026-09-24 14:38:07 -04:00
|
|
|
|
<div id="sheet"></div>
|
|
|
|
|
|
</div>
|
|
|
|
|
|
</main>
|
|
|
|
|
|
|
|
|
|
|
|
<script type="module" src="./js/app.js"></script>
|
|
|
|
|
|
</body>
|
|
|
|
|
|
</html>
|