arthur/index.html

229 lines
14 KiB
HTML
Raw Normal View History

<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>arthur</title>
<style>
:root { --bg:#0b0d13; --panel:#12151f; --line:#232836; --fg:#e8eaf0;
--dim:#8891a5; --accent:#fbbf24; --ok:#4ade80; --err:#f87171; }
* { box-sizing: border-box; }
body { margin:0; background:var(--bg); color:var(--fg);
font:13px/1.5 ui-monospace,SFMono-Regular,Menlo,monospace; }
header { padding:12px 18px; border-bottom:1px solid var(--line);
display:flex; gap:10px; align-items:center; flex-wrap:wrap; }
h1 { font-size:14px; margin:0 8px 0 0; letter-spacing:.04em; }
h1 span { color:var(--dim); font-weight:400; }
main { padding:16px 18px 40px; display:flex; flex-direction:column; gap:16px; }
.row { display:flex; gap:16px; flex-wrap:wrap; }
.panel { background:var(--panel); border:1px solid var(--line); border-radius:4px; padding:12px; }
.panel h2 { font-size:11px; text-transform:uppercase; letter-spacing:.08em;
color:var(--dim); margin:0 0 8px; font-weight:500; }
canvas { display:block; image-rendering:pixelated; max-width:100%; border-radius:2px; }
button,input[type=text],select { font:inherit; background:#1c2130; color:var(--fg);
border:1px solid var(--line); border-radius:3px; padding:5px 10px; cursor:pointer; }
button:hover { border-color:var(--accent); }
input[type=text] { cursor:text; }
label.ctl { display:grid; grid-template-columns:132px 1fr 46px; gap:10px;
align-items:center; margin-bottom:6px; }
label.ctl span:first-child { color:var(--dim); }
label.ctl output { text-align:right; color:var(--accent); }
input[type=range] { width:100%; accent-color:var(--accent); }
#status { color:var(--dim); } #status.ok{color:var(--ok)} #status.err{color:var(--err)} #status.warn{color:var(--accent)}
#readout { color:var(--dim); font-size:12px; }
/* the frame strip is the editing surface */
#strip { display:flex; flex-wrap:wrap; gap:4px; }
.fr { position:relative; cursor:pointer; border:2px solid transparent; border-radius:3px; line-height:0; }
.fr canvas { border-radius:1px; }
.fr span { position:absolute; left:2px; bottom:2px; font-size:10px; line-height:1.2;
padding:0 3px; border-radius:2px; background:#000a; color:#fff; }
.fr.keep { border-color:var(--ok); }
.fr.drop { border-color:#2a2f3e; }
.fr.drop canvas { opacity:.26; filter:grayscale(1); }
.fr.cur { border-color:var(--accent); }
.fr.keep::after { content:'●'; position:absolute; right:3px; top:1px; color:var(--ok); font-size:10px; }
#sheet { display:flex; flex-wrap:wrap; gap:10px; }
#sheet .cell { display:flex; flex-direction:column; gap:3px; cursor:pointer; }
#sheet .cell span { color:var(--dim); font-size:11px; }
#sheet .cell:hover span { color:var(--accent); }
#palette { display:flex; gap:12px; flex-wrap:wrap; }
.sw { display:flex; align-items:center; gap:5px; color:var(--dim); font-size:11px; }
.sw input { width:26px; height:20px; padding:0; border:1px solid var(--line); background:none; }
.legend { color:var(--dim); font-size:11px; margin-top:6px; }
kbd { background:#1c2130; border:1px solid var(--line); border-radius:2px;
padding:0 4px; color:var(--fg); font-size:11px; }
</style>
</head>
<body>
<header>
<h1>arthur <span>— rotoscope + frame removal</span></h1>
<button id="btn-synth">Synthetic</button>
<input type="text" id="framedir" value="frames" size="7" title="frame directory">
<button id="btn-frames">Load frames</button>
<button id="btn-play">Play</button>
Eyes: lids, blinking, line of sight Three parts per eye, stacked the way the mouth is - dark lash ring, sclera inside it, iris inside that, square pupil in the iris. A blink then costs nothing: when the lid shuts the traced ring goes flat and the lash line collapses to a lens, which is a closed eye, drawn correctly, for free. Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a quantised position - and that is where the stylisation lives. Line of sight. Gaze is the iris centre relative to the midpoint of the eye's two corners, in units of corner distance. Both corners are in RIGID, so the origin and the scale are immune to the performance being measured; against the lid ring's centroid instead, every blink would drag the origin down and fake a glance at the floor on exactly the frames where the eye is most visible. Both eyes share one gaze - at this size the difference between the two measurements is noise, not vergence, and independent per-eye noise reads as wall-eyed immediately. Openness stays per-eye so a wink survives. Gaze is then quantised to a pixel grid with a dwell, which is not a stylisation imposed on the truth: real eyes move in saccades, and the smooth drift left in the measurement is tracker noise plus head-compensation error. Snapping to a grid removes the noise and recovers the saccade in one operation. The iris is placed in the frame of the already-smoothed, already-subsampled lid ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any even budget keeps them at 0 and n/2 - so it cannot drift relative to its own eye. Size is authored from the take mean, never remeasured per frame: a radius that breathes by a fraction of a pixel flickers a pixel on and off around the whole silhouette. iris anchor toggles steady/free/locked, because how much the eye wanders turns out to be an aesthetic choice and not only a correctness one. Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not have: blink hold. A blink is one frame at 12fps and a single frame of closed eye reads as a dropped frame, so once the eye shuts it stays shut long enough to be legible. Detection accuracy is not the problem; legibility is. The pupil is a square because at three pixels a circle is a plus sign with the corners gnawed off, and it changes shape as it moves. Drawn from a rounded centre shared with the iris so it is exactly its nominal size on every frame. Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro would: the lid crops the iris at extreme gaze for free, so nothing has to clamp the gaze, which would flatten the performance at the extremes that carry it. Which iris block belongs to which eye is RESOLVED from geometry, not declared. A swap looks almost right - each eye still has a disc roughly where it belongs - so it survives an eyeball and then reads as a subtly wall-eyed character forever. Voted across every frame; the test feeds a deliberately swapped track. Also: exposure. Aesthetic sparseness was set by the extraction rate, which made the timing a property of a directory of PNGs - auditioning 12 against 24 meant re-ripping and re-detecting the whole clip. It is now a render-time grid, on 1s/2s/3s/4s, so the dense track keeps everything and the audio clock is untouched. The take format already carried an exposure field; it was never driven. Everything rides the same grid, because a head cutting on the odd frames while the mouth cuts on the even ones reads as two performances laid over each other. 41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a swapped track, a blink does not fake a change of gaze, a stencilled disc cannot spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure never reads a pose from the future. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
<select id="irisAnchor" title="what the iris hangs off">
<option value="steady" selected>iris: steady</option>
<option value="free">iris: free</option>
<option value="locked">iris: locked</option>
</select>
<select id="gazeOrigin" title="what counts as looking straight ahead">
<option value="median" selected>origin: median</option>
<option value="neutral">origin: neutral f</option>
</select>
<select id="exposure" title="exposure — how often the picture gets a new drawing">
<option value="1" selected>on 1s</option>
<option value="2">on 2s</option>
<option value="3">on 3s</option>
<option value="4">on 4s</option>
</select>
<select id="speed" title="playback speed"><option value="1">1x</option><option value="0.5">½x</option><option value="0.25">¼x</option></select>
<audio id="audio" controls hidden style="height:28px;vertical-align:middle"></audio>
<button id="btn-keepall">Keep all</button>
<button id="btn-suggest">Suggest</button>
<input type="text" id="takename" value="line_01" size="9" title="take name">
<select id="plateMode" title="plate representation (B to cycle)">
<option value="photo-dim">photo dim</option>
<option value="photo">photo</option>
<option value="posterize">posterized</option>
<option value="oval" selected>oval</option>
<option value="oval+photo">oval + photo</option>
<option value="off">none</option>
</select>
<button id="btn-saveframe">Save frame 4x</button>
<button id="btn-export">Export .take</button>
<span id="status"></span>
</header>
<main>
<div class="row">
<div class="panel">
<h2>source + landmarks</h2>
<canvas id="cv-source"></canvas>
Eyes: lids, blinking, line of sight Three parts per eye, stacked the way the mouth is - dark lash ring, sclera inside it, iris inside that, square pupil in the iris. A blink then costs nothing: when the lid shuts the traced ring goes flat and the lash line collapses to a lens, which is a closed eye, drawn correctly, for free. Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a quantised position - and that is where the stylisation lives. Line of sight. Gaze is the iris centre relative to the midpoint of the eye's two corners, in units of corner distance. Both corners are in RIGID, so the origin and the scale are immune to the performance being measured; against the lid ring's centroid instead, every blink would drag the origin down and fake a glance at the floor on exactly the frames where the eye is most visible. Both eyes share one gaze - at this size the difference between the two measurements is noise, not vergence, and independent per-eye noise reads as wall-eyed immediately. Openness stays per-eye so a wink survives. Gaze is then quantised to a pixel grid with a dwell, which is not a stylisation imposed on the truth: real eyes move in saccades, and the smooth drift left in the measurement is tracker noise plus head-compensation error. Snapping to a grid removes the noise and recovers the saccade in one operation. The iris is placed in the frame of the already-smoothed, already-subsampled lid ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any even budget keeps them at 0 and n/2 - so it cannot drift relative to its own eye. Size is authored from the take mean, never remeasured per frame: a radius that breathes by a fraction of a pixel flickers a pixel on and off around the whole silhouette. iris anchor toggles steady/free/locked, because how much the eye wanders turns out to be an aesthetic choice and not only a correctness one. Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not have: blink hold. A blink is one frame at 12fps and a single frame of closed eye reads as a dropped frame, so once the eye shuts it stays shut long enough to be legible. Detection accuracy is not the problem; legibility is. The pupil is a square because at three pixels a circle is a plus sign with the corners gnawed off, and it changes shape as it moves. Drawn from a rounded centre shared with the iris so it is exactly its nominal size on every frame. Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro would: the lid crops the iris at extreme gaze for free, so nothing has to clamp the gaze, which would flatten the performance at the extremes that carry it. Which iris block belongs to which eye is RESOLVED from geometry, not declared. A swap looks almost right - each eye still has a disc roughly where it belongs - so it survives an eyeball and then reads as a subtly wall-eyed character forever. Voted across every frame; the test feeds a deliberately swapped track. Also: exposure. Aesthetic sparseness was set by the extraction rate, which made the timing a property of a directory of PNGs - auditioning 12 against 24 meant re-ripping and re-detecting the whole clip. It is now a render-time grid, on 1s/2s/3s/4s, so the dense track keeps everything and the audio clock is untouched. The take format already carried an exposure field; it was never driven. Everything rides the same grid, because a head cutting on the odd frames while the mouth cuts on the even ones reads as two performances laid over each other. 41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a swapped track, a blink does not fake a change of gaze, a stencilled disc cannot spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure never reads a pose from the future. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
<div class="legend">outer lip <b style="color:#4ade80">—</b> · inner lip <b style="color:#f87171">—</b> ·
lids <b style="color:#60a5fa">—</b> · iris <b style="color:#fbbf24">—</b></div>
</div>
<div class="panel">
<h2>stabilised (head-local)</h2>
<canvas id="cv-stab"></canvas>
Eyes: lids, blinking, line of sight Three parts per eye, stacked the way the mouth is - dark lash ring, sclera inside it, iris inside that, square pupil in the iris. A blink then costs nothing: when the lid shuts the traced ring goes flat and the lash line collapses to a lens, which is a closed eye, drawn correctly, for free. Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a quantised position - and that is where the stylisation lives. Line of sight. Gaze is the iris centre relative to the midpoint of the eye's two corners, in units of corner distance. Both corners are in RIGID, so the origin and the scale are immune to the performance being measured; against the lid ring's centroid instead, every blink would drag the origin down and fake a glance at the floor on exactly the frames where the eye is most visible. Both eyes share one gaze - at this size the difference between the two measurements is noise, not vergence, and independent per-eye noise reads as wall-eyed immediately. Openness stays per-eye so a wink survives. Gaze is then quantised to a pixel grid with a dwell, which is not a stylisation imposed on the truth: real eyes move in saccades, and the smooth drift left in the measurement is tracker noise plus head-compensation error. Snapping to a grid removes the noise and recovers the saccade in one operation. The iris is placed in the frame of the already-smoothed, already-subsampled lid ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any even budget keeps them at 0 and n/2 - so it cannot drift relative to its own eye. Size is authored from the take mean, never remeasured per frame: a radius that breathes by a fraction of a pixel flickers a pixel on and off around the whole silhouette. iris anchor toggles steady/free/locked, because how much the eye wanders turns out to be an aesthetic choice and not only a correctness one. Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not have: blink hold. A blink is one frame at 12fps and a single frame of closed eye reads as a dropped frame, so once the eye shuts it stays shut long enough to be legible. Detection accuracy is not the problem; legibility is. The pupil is a square because at three pixels a circle is a plus sign with the corners gnawed off, and it changes shape as it moves. Drawn from a rounded centre shared with the iris so it is exactly its nominal size on every frame. Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro would: the lid crops the iris at extreme gaze for free, so nothing has to clamp the gaze, which would flatten the performance at the extremes that carry it. Which iris block belongs to which eye is RESOLVED from geometry, not declared. A swap looks almost right - each eye still has a disc roughly where it belongs - so it survives an eyeball and then reads as a subtly wall-eyed character forever. Voted across every frame; the test feeds a deliberately swapped track. Also: exposure. Aesthetic sparseness was set by the extraction rate, which made the timing a property of a directory of PNGs - auditioning 12 against 24 meant re-ripping and re-detecting the whole clip. It is now a render-time grid, on 1s/2s/3s/4s, so the dense track keeps everything and the audio clock is untouched. The take format already carried an exposure field; it was never driven. Everything rides the same grid, because a head cutting on the odd frames while the mouth cuts on the even ones reads as two performances laid over each other. 41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a swapped track, a blink does not fake a change of gaze, a stencilled disc cannot spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure never reads a pose from the future. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
<div class="legend">should sit still except the mouth and eyes · grey = held plate outline<br>
dark green ghost = unshifted mouth when lead ≠ 0 · red lid ring = blink</div>
</div>
<div class="panel">
<h2>flat render — 320×200 indexed</h2>
<canvas id="cv-render"></canvas>
<div class="legend" id="framelabel"></div>
<div class="legend">audio drives the clock — dropped frames, never drift<br>
Eyes: lids, blinking, line of sight Three parts per eye, stacked the way the mouth is - dark lash ring, sclera inside it, iris inside that, square pupil in the iris. A blink then costs nothing: when the lid shuts the traced ring goes flat and the lash line collapses to a lens, which is a closed eye, drawn correctly, for free. Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a quantised position - and that is where the stylisation lives. Line of sight. Gaze is the iris centre relative to the midpoint of the eye's two corners, in units of corner distance. Both corners are in RIGID, so the origin and the scale are immune to the performance being measured; against the lid ring's centroid instead, every blink would drag the origin down and fake a glance at the floor on exactly the frames where the eye is most visible. Both eyes share one gaze - at this size the difference between the two measurements is noise, not vergence, and independent per-eye noise reads as wall-eyed immediately. Openness stays per-eye so a wink survives. Gaze is then quantised to a pixel grid with a dwell, which is not a stylisation imposed on the truth: real eyes move in saccades, and the smooth drift left in the measurement is tracker noise plus head-compensation error. Snapping to a grid removes the noise and recovers the saccade in one operation. The iris is placed in the frame of the already-smoothed, already-subsampled lid ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any even budget keeps them at 0 and n/2 - so it cannot drift relative to its own eye. Size is authored from the take mean, never remeasured per frame: a radius that breathes by a fraction of a pixel flickers a pixel on and off around the whole silhouette. iris anchor toggles steady/free/locked, because how much the eye wanders turns out to be an aesthetic choice and not only a correctness one. Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not have: blink hold. A blink is one frame at 12fps and a single frame of closed eye reads as a dropped frame, so once the eye shuts it stays shut long enough to be legible. Detection accuracy is not the problem; legibility is. The pupil is a square because at three pixels a circle is a plus sign with the corners gnawed off, and it changes shape as it moves. Drawn from a rounded centre shared with the iris so it is exactly its nominal size on every frame. Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro would: the lid crops the iris at extreme gaze for free, so nothing has to clamp the gaze, which would flatten the performance at the extremes that carry it. Which iris block belongs to which eye is RESOLVED from geometry, not declared. A swap looks almost right - each eye still has a disc roughly where it belongs - so it survives an eyeball and then reads as a subtly wall-eyed character forever. Voted across every frame; the test feeds a deliberately swapped track. Also: exposure. Aesthetic sparseness was set by the extraction rate, which made the timing a property of a directory of PNGs - auditioning 12 against 24 meant re-ripping and re-detecting the whole clip. It is now a render-time grid, on 1s/2s/3s/4s, so the dense track keeps everything and the audio clock is untouched. The take format already carried an exposure field; it was never driven. Everything rides the same grid, because a head cutting on the odd frames while the mouth cuts on the even ones reads as two performances laid over each other. 41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a swapped track, a blink does not fake a change of gaze, a stencilled disc cannot spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure never reads a pose from the future. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
<b>exposure</b> holds the picture on a grid: rip at 24, render on 2s for
12. The dense track and the audio are untouched, so it is reversible<br>
plate representation: <kbd>B</kbd> cycles · photo modes are <b>registered</b>
into raster space, so tracing them lands on the mouth</div>
</div>
</div>
<div class="panel">
<h2>frames — green kept (gets its own drawing) · dim held from the last kept frame</h2>
<div id="strip"></div>
<input type="range" id="scrub" min="0" max="0" value="0" style="width:100%;margin-top:10px">
<div class="legend">
<kbd>←</kbd> <kbd>→</kbd> step · <kbd>X</kbd> delete · <kbd>K</kbd> keep · <kbd>B</kbd> background ·
<kbd>[</kbd> <kbd>]</kbd> lead ·
click to select · double-click or shift-click to toggle ·
the mouth keeps <b>every</b> frame regardless
</div>
<div class="legend" id="readout"></div>
</div>
<div class="row">
<div class="panel" style="flex:1 1 400px">
<h2>knobs</h2>
<label class="ctl"><span>vertices</span><input type="range" id="verts" min="4" max="16" step="2" value="8"><output id="vertsv"></output></label>
<label class="ctl"><span>mouth lead ±f</span><input type="range" id="lead" min="-6" max="6" value="0"><output id="leadv"></output></label>
<label class="ctl"><span>contour avg ±f</span><input type="range" id="contourSmooth" min="0" max="4" value="1"><output id="contourSmoothv"></output></label>
<label class="ctl"><span>anchor avg ±f</span><input type="range" id="smoothWin" min="0" max="8" value="2"><output id="smoothWinv"></output></label>
<label class="ctl"><span>closed-mouth cut</span><input type="range" id="apertureThresh" min="0" max="400" value="120"><output id="apertureThreshv"></output></label>
<label class="ctl"><span>teeth contrast</span><input type="range" id="teethOn" min="1" max="60" value="16"><output id="teethOnv"></output></label>
<label class="ctl"><span>cavity erode</span><input type="range" id="teethErode" min="0" max="45" value="18"><output id="teethErodev"></output></label>
<label class="ctl"><span>blob grow/erode</span><input type="range" id="blobGrow" min="-4" max="4" value="0"><output id="blobGrowv"></output></label>
<label class="ctl"><span>tongue reject</span><input type="range" id="tongueReject" min="2" max="40" value="18"><output id="tongueRejectv"></output></label>
<label class="ctl"><span>prefer upper</span><input type="range" id="topBias" min="0" max="120" value="60"><output id="topBiasv"></output></label>
<label class="ctl"><span>teeth vertices</span><input type="range" id="teethVerts" min="5" max="20" value="10"><output id="teethVertsv"></output></label>
<label class="ctl"><span>teeth avg ±f</span><input type="range" id="teethSmooth" min="0" max="4" value="1"><output id="teethSmoothv"></output></label>
<label class="ctl"><span>teeth dwell</span><input type="range" id="teethDwell" min="0" max="6" value="1"><output id="teethDwellv"></output></label>
Eyes: lids, blinking, line of sight Three parts per eye, stacked the way the mouth is - dark lash ring, sclera inside it, iris inside that, square pupil in the iris. A blink then costs nothing: when the lid shuts the traced ring goes flat and the lash line collapses to a lens, which is a closed eye, drawn correctly, for free. Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a quantised position - and that is where the stylisation lives. Line of sight. Gaze is the iris centre relative to the midpoint of the eye's two corners, in units of corner distance. Both corners are in RIGID, so the origin and the scale are immune to the performance being measured; against the lid ring's centroid instead, every blink would drag the origin down and fake a glance at the floor on exactly the frames where the eye is most visible. Both eyes share one gaze - at this size the difference between the two measurements is noise, not vergence, and independent per-eye noise reads as wall-eyed immediately. Openness stays per-eye so a wink survives. Gaze is then quantised to a pixel grid with a dwell, which is not a stylisation imposed on the truth: real eyes move in saccades, and the smooth drift left in the measurement is tracker noise plus head-compensation error. Snapping to a grid removes the noise and recovers the saccade in one operation. The iris is placed in the frame of the already-smoothed, already-subsampled lid ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any even budget keeps them at 0 and n/2 - so it cannot drift relative to its own eye. Size is authored from the take mean, never remeasured per frame: a radius that breathes by a fraction of a pixel flickers a pixel on and off around the whole silhouette. iris anchor toggles steady/free/locked, because how much the eye wanders turns out to be an aesthetic choice and not only a correctness one. Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not have: blink hold. A blink is one frame at 12fps and a single frame of closed eye reads as a dropped frame, so once the eye shuts it stays shut long enough to be legible. Detection accuracy is not the problem; legibility is. The pupil is a square because at three pixels a circle is a plus sign with the corners gnawed off, and it changes shape as it moves. Drawn from a rounded centre shared with the iris so it is exactly its nominal size on every frame. Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro would: the lid crops the iris at extreme gaze for free, so nothing has to clamp the gaze, which would flatten the performance at the extremes that carry it. Which iris block belongs to which eye is RESOLVED from geometry, not declared. A swap looks almost right - each eye still has a disc roughly where it belongs - so it survives an eyeball and then reads as a subtly wall-eyed character forever. Voted across every frame; the test feeds a deliberately swapped track. Also: exposure. Aesthetic sparseness was set by the extraction rate, which made the timing a property of a directory of PNGs - auditioning 12 against 24 meant re-ripping and re-detecting the whole clip. It is now a render-time grid, on 1s/2s/3s/4s, so the dense track keeps everything and the audio clock is untouched. The take format already carried an exposure field; it was never driven. Everything rides the same grid, because a head cutting on the odd frames while the mouth cuts on the even ones reads as two performances laid over each other. 41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a swapped track, a blink does not fake a change of gaze, a stencilled disc cannot spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure never reads a pose from the future. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
<label class="ctl"><span>eye vertices</span><input type="range" id="eyeVerts" min="4" max="12" step="2" value="8"><output id="eyeVertsv"></output></label>
<label class="ctl"><span>lash line</span><input type="range" id="lashPx" min="0" max="3" value="1"><output id="lashPxv"></output></label>
<label class="ctl"><span>blink cut</span><input type="range" id="blinkCut" min="20" max="300" value="130"><output id="blinkCutv"></output></label>
<label class="ctl"><span>blink hold ±f</span><input type="range" id="blinkHold" min="1" max="5" value="2"><output id="blinkHoldv"></output></label>
<label class="ctl"><span>blink dwell</span><input type="range" id="blinkDwell" min="0" max="4" value="0"><output id="blinkDwellv"></output></label>
<label class="ctl"><span>gaze gain</span><input type="range" id="gazeGain" min="50" max="400" value="100"><output id="gazeGainv"></output></label>
<label class="ctl"><span>gaze step</span><input type="range" id="gazeStep" min="0" max="6" value="2"><output id="gazeStepv"></output></label>
<label class="ctl"><span>gaze dwell</span><input type="range" id="gazeDwell" min="0" max="6" value="2"><output id="gazeDwellv"></output></label>
<label class="ctl"><span>iris size</span><input type="range" id="irisSize" min="20" max="70" value="42"><output id="irisSizev"></output></label>
<label class="ctl"><span>pupil</span><input type="range" id="pupilPx" min="0" max="7" value="3"><output id="pupilPxv"></output></label>
<label class="ctl"><span>suggest tolerance</span><input type="range" id="tol" min="2" max="60" value="14"><output id="tolv"></output></label>
<div class="legend">
<b>mouth lead</b> shifts the performance tracks earlier (positive) against
the audio and the head. Averaging has no phase lag but it blurs onsets, so
an opening reads later than it is; animators also draw mouths a frame or
two ahead of the sound as standard practice. <kbd>[</kbd> <kbd>]</kbd>.<br>
<b>contour avg</b> 0 = off, 1 = ±1 frame. Removes per-frame landmark
jitter. Push past 2 and it starts eating articulation.<br>
<b>anchor avg</b> smooths the head transform only — never the contour.<br>
<b>teeth contrast</b> gates on how far apart the cavity's dark and bright
halves are — Otsu always returns <i>some</i> threshold, so this is what
stops it inventing teeth in a dark mouth.
<b>cavity erode</b> pulls the sampled region in from the lip edge;
<b>blob grow/erode</b> resizes the found blob itself.
<b>tongue reject</b> drops pixels that are red relative to their own
brightness; <b>prefer upper</b> biases component choice toward the top of
the cavity, where teeth are and the tongue is not.
<b>dwell</b> is how many frames a presence change must persist.<br>
Eyes: lids, blinking, line of sight Three parts per eye, stacked the way the mouth is - dark lash ring, sclera inside it, iris inside that, square pupil in the iris. A blink then costs nothing: when the lid shuts the traced ring goes flat and the lash line collapses to a lens, which is a closed eye, drawn correctly, for free. Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a quantised position - and that is where the stylisation lives. Line of sight. Gaze is the iris centre relative to the midpoint of the eye's two corners, in units of corner distance. Both corners are in RIGID, so the origin and the scale are immune to the performance being measured; against the lid ring's centroid instead, every blink would drag the origin down and fake a glance at the floor on exactly the frames where the eye is most visible. Both eyes share one gaze - at this size the difference between the two measurements is noise, not vergence, and independent per-eye noise reads as wall-eyed immediately. Openness stays per-eye so a wink survives. Gaze is then quantised to a pixel grid with a dwell, which is not a stylisation imposed on the truth: real eyes move in saccades, and the smooth drift left in the measurement is tracker noise plus head-compensation error. Snapping to a grid removes the noise and recovers the saccade in one operation. The iris is placed in the frame of the already-smoothed, already-subsampled lid ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any even budget keeps them at 0 and n/2 - so it cannot drift relative to its own eye. Size is authored from the take mean, never remeasured per frame: a radius that breathes by a fraction of a pixel flickers a pixel on and off around the whole silhouette. iris anchor toggles steady/free/locked, because how much the eye wanders turns out to be an aesthetic choice and not only a correctness one. Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not have: blink hold. A blink is one frame at 12fps and a single frame of closed eye reads as a dropped frame, so once the eye shuts it stays shut long enough to be legible. Detection accuracy is not the problem; legibility is. The pupil is a square because at three pixels a circle is a plus sign with the corners gnawed off, and it changes shape as it moves. Drawn from a rounded centre shared with the iris so it is exactly its nominal size on every frame. Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro would: the lid crops the iris at extreme gaze for free, so nothing has to clamp the gaze, which would flatten the performance at the extremes that carry it. Which iris block belongs to which eye is RESOLVED from geometry, not declared. A swap looks almost right - each eye still has a disc roughly where it belongs - so it survives an eyeball and then reads as a subtly wall-eyed character forever. Voted across every frame; the test feeds a deliberately swapped track. Also: exposure. Aesthetic sparseness was set by the extraction rate, which made the timing a property of a directory of PNGs - auditioning 12 against 24 meant re-ripping and re-detecting the whole clip. It is now a render-time grid, on 1s/2s/3s/4s, so the dense track keeps everything and the audio clock is untouched. The take format already carried an exposure field; it was never driven. Everything rides the same grid, because a head cutting on the odd frames while the mouth cuts on the even ones reads as two performances laid over each other. 41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a swapped track, a blink does not fake a change of gaze, a stencilled disc cannot spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure never reads a pose from the future. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
<b>blink cut</b> is lid gap over corner distance — normalised, so one
value carries across takes. <b>blink hold</b> is the minimum length of a
blink: a real blink is one frame at 12fps and a single frame of closed
eye reads as a dropout, so it is extended to a beat.
<b>pupil</b> is a square, in whole pixels, 0 to turn it off: at this size
a circle of radius 1.5 is a plus sign with the corners gnawed off and it
changes shape as it moves, where a square stays the mark you drew.<br>
<b>gaze step</b> is the grid the iris snaps to, in raster pixels, and
<b>gaze dwell</b> is how long a new cell must hold — together they turn
drift into saccades. <b>gaze gain</b> exaggerates or damps the throw;
measured excursion is small and a character usually wants more of it.<br>
<b>suggest tolerance</b> only affects the Suggest button: max head movement
allowed before a new drawing is required.
</div>
</div>
Eyes: lids, blinking, line of sight Three parts per eye, stacked the way the mouth is - dark lash ring, sclera inside it, iris inside that, square pupil in the iris. A blink then costs nothing: when the lid shuts the traced ring goes flat and the lash line collapses to a lens, which is a closed eye, drawn correctly, for free. Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a quantised position - and that is where the stylisation lives. Line of sight. Gaze is the iris centre relative to the midpoint of the eye's two corners, in units of corner distance. Both corners are in RIGID, so the origin and the scale are immune to the performance being measured; against the lid ring's centroid instead, every blink would drag the origin down and fake a glance at the floor on exactly the frames where the eye is most visible. Both eyes share one gaze - at this size the difference between the two measurements is noise, not vergence, and independent per-eye noise reads as wall-eyed immediately. Openness stays per-eye so a wink survives. Gaze is then quantised to a pixel grid with a dwell, which is not a stylisation imposed on the truth: real eyes move in saccades, and the smooth drift left in the measurement is tracker noise plus head-compensation error. Snapping to a grid removes the noise and recovers the saccade in one operation. The iris is placed in the frame of the already-smoothed, already-subsampled lid ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any even budget keeps them at 0 and n/2 - so it cannot drift relative to its own eye. Size is authored from the take mean, never remeasured per frame: a radius that breathes by a fraction of a pixel flickers a pixel on and off around the whole silhouette. iris anchor toggles steady/free/locked, because how much the eye wanders turns out to be an aesthetic choice and not only a correctness one. Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not have: blink hold. A blink is one frame at 12fps and a single frame of closed eye reads as a dropped frame, so once the eye shuts it stays shut long enough to be legible. Detection accuracy is not the problem; legibility is. The pupil is a square because at three pixels a circle is a plus sign with the corners gnawed off, and it changes shape as it moves. Drawn from a rounded centre shared with the iris so it is exactly its nominal size on every frame. Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro would: the lid crops the iris at extreme gaze for free, so nothing has to clamp the gaze, which would flatten the performance at the extremes that carry it. Which iris block belongs to which eye is RESOLVED from geometry, not declared. A swap looks almost right - each eye still has a disc roughly where it belongs - so it survives an eyeball and then reads as a subtly wall-eyed character forever. Voted across every frame; the test feeds a deliberately swapped track. Also: exposure. Aesthetic sparseness was set by the extraction rate, which made the timing a property of a directory of PNGs - auditioning 12 against 24 meant re-ripping and re-detecting the whole clip. It is now a render-time grid, on 1s/2s/3s/4s, so the dense track keeps everything and the audio clock is untouched. The take format already carried an exposure field; it was never driven. Everything rides the same grid, because a head cutting on the odd frames while the mouth cuts on the even ones reads as two performances laid over each other. 41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a swapped track, a blink does not fake a change of gaze, a stencilled disc cannot spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure never reads a pose from the future. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-24 18:06:04 -04:00
<div class="panel" style="flex:0 1 170px">
<h2>gaze field</h2>
<canvas id="cv-gaze"></canvas>
<div class="legend" id="eyeinfo" style="white-space:pre-line"></div>
<div class="legend">green = every cell the iris visits in the take ·
grey = raw · amber = where it is now, quantised</div>
</div>
<div class="panel" style="flex:0 1 190px">
<h2>teeth measurement</h2>
<div id="cv-teeth"></div>
<div class="legend" id="teethinfo"></div>
<div class="legend">green = kept pixels · amber = extracted contour</div>
</div>
<div class="panel" style="flex:1 1 240px">
<h2>palette</h2>
<div id="palette"></div>
<div class="legend" style="margin-top:12px">
Flat indexed fills, no antialiasing — the rasteriser writes palette
indices, the way <code>csd_render_poly</code> does.
</div>
</div>
</div>
<div class="panel">
<h2>drawings needed — registered reference per kept frame, with the range it holds</h2>
<div id="sheet"></div>
</div>
</main>
<script type="module" src="./js/app.js"></script>
</body>
</html>