- scene: mark-group model (resolve/selection->marks/seg-local/marks->rows); drop legacy tl.marks/tl.md - authoring: draft is a normal annotation group flagged :draft; generic put-group + with-let cancel; two-input click-to-fill editor in mark time - timeline: dashed draft preview, numeric track ordering - video frame readout (track+clip, mark time) clickable to fill an endpoint - discontinuity jump popover; nested-annotation count badge - localStorage persistence of the authored annotation layer Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
10 KiB
Data Model
a mark-group is the fundamental datastructure of this app. the whole scene graph is composed of them. we use this same structure to represent the main timeline, the clips within the timeline, and the annotations on the timeline. a mark-group is just an ordered list of marks with a little bit of extra data hanging off them depending on the type. it also has a parent: the mark-group context it belongs to.
a mark can represent an instant or a range of time, and that instant or range can be defined in terms of frames (relative to the current timeline on the timeline stack) or frames w/r/t to other mark-groups and these can be mixed and matched. each mark can also optionally specify a target video track (video tracks and clips are initially sourced from an initial otio file; the otio is only a seed – it populates :tracks and the initial clip + timeline mark-groups once, then we never look at it again. :tracks is the only thing in the whole app that isn't a mark-group). thus, the main timeline is a mark group with one mark: the start and end timestamp. a clip which belongs to that main timeline is a mark group with one mark: the start and end timestamps (within the main timeline) AND a video track. but the track hangs off the mark, optionally, not the root of the clip. a clip is only different from an annotation in that one of its marks specifies a track (and, soon, :thumbnails), so an annotation could target a track too. clips are parentless: they're a flat pool, referenced by id, never owned. :parent is an annotation-only thing – the authoring/visibility context, i.e. which timeline i was in when i made it. this dodges the whole knot: if clips had parents, expanding an annotation would have to make the clip a child of the annotation AND the main timeline AND every sub-annotation at once. containment is a reference, not ownership. an annotation is a mark group which conceptually represents a point or points of interest with optional commentary, though it looks not substantially different than a timeline or clip in the data structure, and really it's so flexible it could represent whole re-edits of clips. its marks can also be timestamps in the context of some timeline, or they can be the start and end frames of a given clip, defined relative to the clip, or defined relative to another annotation, or any combination thereof, in any amount, and with any combination of ranges and instants. an instant is just a range where start == end (length 0), no separate type. concatenation goes by length, so an instant adds no duration – it's a marker at the current offset, drawn as a diamond instead of a bar.
because an annotation is just a mark-group, and a timeline is just a mark group, any annotation can be pushed onto the timeline-stack, replacing the main timeline. the clips within the marks in the mark group are laid end to end to form one continuous duration, a new timeline. the cool thing here is that you can now annotate within the context of this annotation. so if we are inside annotation A, annotation A is the :parent of our new Annotation B. if annotation B uses absolute timestamp marks, they are relative to the annotation A timeline, not the main timeline. and if annotation B uses clip based timestamps, they can only reference the frames of the underlying clip which are within range of the annotation (annotation A, the parent, may have start half way through the clip at the beginning, and end half way through the clip at the end). and then you can push annotation B onto the timeline stack, annotate within that, and on and on.
something about mark groups to note is that marks need not be defined in order. if the main timeline has clip A, B, and C laid end to end, an annotation X can have marks /noe/tl/src/commit/a62af2584bd688ca9febe9a548a900c33e47c2b6/tl/clipC%5B0%5D,%20clipC%5B-1, [clipA[0], clipA[-1]], [clipB[0], clipB[-1]]. -1 represents last available frame of the clip (note here that available frame may differ from absolute last frame of the underlying clip, because the annotation context we're in could cut off half the clip, for example). note this clamping only bites for raw clip refs across a trim – if you reference the subclip (the parent's mark) instead, the trim is baked into the mark's range, so subclip[-1] = the subclip's own end, no clamp (more below). in this example, we have totally rearranged the clips into a timeline B C A, end to end. and you can also imagine we can cut clips in half, interleave them, repeat them and so on.
hmm here's a struggle though. let's say i have interleaved half of A with half of B in an annotation which i pushed onto the timeline stack:
marks: /noe/tl/src/commit/a62af2584bd688ca9febe9a548a900c33e47c2b6/tl/clipB%5B20%5D,%20clipB%5B40, [clipA[0], clipA[20]], [clipB[0], clipB[20]] [clipA[20], clipA[40]]]
we want to be able to mark these sub clips independently for a new annotation Y, right? but these are just 2 root clips that became 4. so how can we define annotation Y with respect to any of these 4 clips? we don't want to use absolute timeline time, but we also don't want to use absolute clip time because our marks can cross between the first sub clip (b 20 - 40) and the second (a 0 to a 20), so clip time means nothing here. we also need to remember that the parent timeline can clear out its marks at will. so we can definitely orphan annotations - that's ok, that's a UI concern we can display warnings for and just grey out basically (drop to bottom of annotation list, for example, with a warn emoji and let the user edit to specify its marks again, even warn on which marks are broken, and if at least one mark is still valid still display it there). so it rly seems like an annotation SHOULD create synthetic subclips which point at the raw clips so that we can use the synthetic subclips as our targets. but they need to be stable identities, serializable/deserializable.
resolved: every mark gets a stable id (a uuid) at creation, and a clip/subclip-ref mark stays within ONE clip – so the addressable subclip is just a single-clip mark, addressed by mark-id alone. a selection that crosses clip boundaries is stored as a RUN of per-clip marks (the lane sticks the contiguous run into one bar; the editor splits/merges at boundaries on edit). a mark is already a recursive structure pointing at a raw clip, the id just makes it addressable, so no separate entities, no recreation lifecycle. Y references A's subclips by mark-id: {:ref <mark-id> :at n}.
- edit a mark's range -> same id -> Y follows it
- reorder marks -> ids travel with them -> Y follows the content, not the slot
- delete a mark -> id gone -> Y dangles -> orphan (grey out, warn per broken mark, keep if at least one still resolves)
- add a mark -> new id
so "keep identity unless the whole thing is different" isn't an algorithm we run, it's a ui affordance: editing a row in place keeps the id, delete-row + add-row makes a new id. like keyed list editing / db rows with primary keys – you carry stable keys, you never diff structure to guess identity.
and the subclip bakes the trim into its own definition, so it's a clean map to source: subclip[n] = src-start + n, subclip[-1] = src-end, no clamp, no context lookup. that makes resolving a ref context-free: (mark-id, scene) -> source range, no timeline-stack needed. the stack only decides which context's local timeline you're looking at, not how a ref resolves. only absolute (bare number) points are context-dependent – they're local frames of the context the mark lives in.
so the whole mark grammar is two point kinds:
- ref point {:ref <mark-id> :at n} -> frame n of that mark's resolved range (context-free, n negative = from the end)
- absolute point <number> -> a local frame of the context the mark lives in (resolved through that context's spans)
mix them in a single range, instant = start==end, track optional on the mark. a raw clip is just a mark whose range is full source + a track. a clip/subclip-ref mark keeps both endpoints on the SAME target, so it resolves to exactly one source segment; only an absolute mark may span several.
decided: a clip/subclip-ref mark never crosses a clip boundary, so the "range whose endpoints are in different subclips" case just can't happen. a selection across clips is a run of single-clip marks instead, one per clip, laid end to end (genuinely contiguous – the lane only draws them as one bar). this keeps resolve trivial (one source segment per ref mark), makes every piece addressable by mark-id alone, and makes a reorder follow each piece independently instead of swelling. the merge/unmerge is localized: dragging a boundary WITHIN a clip edits that end mark in place (id stable); dragging ACROSS a clip boundary adds/removes a whole clip from the selection (creates/destroys that end mark); the fully-contained middle clips never churn. the same rule applies at any depth – "clip boundary" means a boundary in the fully-resolved footage, so it works the same whether you're at root crossing raw clips or nested crossing a parent mark's segments. the one multi-segment mark left is an absolute one (bare-number local range): it's arrangement-relative, you build it by dragging the timeline rather than by referencing, and it resolves via slice.
this brings us to playing. since right now there is only one source video file, we need to be able to seek to arbitrary frames. each context (mark-group) keeps its own local playhead, used when it's the top of the timeline stack. when we hit play, in the example above of annotation X, we find the clip under the playhead, compute the source frame, seek there, and start playing. one correction though: the local playhead has to be the master clock, not the video. you can't derive local position from currentTime – once an annotation repeats or reorders clips, one source frame maps to several local frames, it's not invertible. so the local playhead advances on its own (wall-clock x fps while playing), and every frame we compute expected = group->media(local) and seek the video there only if round(currentTime*fps) != expected. within a clip, expected tracks the video's natural playback so no seek fires; at a mark boundary it jumps once and we seek. and right – no recursion at play time: we resolve the current context once into flat ordered spans, and group->media is just the flat lookup the renderer already does.