14 KiB
Data Model
a mark-group is the fundamental datastructure of this app. the whole scene graph is composed of them. we use this same structure to represent the main timeline, the clips within the timeline, and the annotations on the timeline. a mark-group is just an ordered list of marks with a little bit of extra data hanging off them depending on the type. it also has a parent: the mark-group context it belongs to.
a mark can represent an instant or a range of time, and that instant or range can be defined in terms of frames (relative to the current timeline on the timeline stack) or frames w/r/t to other mark-groups and these can be mixed and matched. each mark can also optionally specify a target video track (video tracks and clips are initially sourced from an initial otio file; the otio is only a seed – it populates :tracks and the initial clip + timeline mark-groups once, then we never look at it again. :tracks is the only thing in the whole app that isn't a mark-group). thus, the main timeline is a mark group with one mark: the start and end timestamp. a clip which belongs to that main timeline is a mark group with one mark: the start and end timestamps (within the main timeline) AND a video track. but the track hangs off the mark, optionally, not the root of the clip. a clip is only different from an annotation in that one of its marks specifies a track (and, soon, :thumbnails), so an annotation could target a track too. clips are parentless: they're a flat pool, referenced by id, never owned. :parent is an annotation-only thing – the authoring/visibility context, i.e. which timeline i was in when i made it. this dodges the whole knot: if clips had parents, expanding an annotation would have to make the clip a child of the annotation AND the main timeline AND every sub-annotation at once. containment is a reference, not ownership. an annotation is a mark group which conceptually represents a point or points of interest with optional commentary, though it looks not substantially different than a timeline or clip in the data structure, and really it's so flexible it could represent whole re-edits of clips. its marks can also be timestamps in the context of some timeline, or they can be the start and end frames of a given clip, defined relative to the clip, or defined relative to another annotation, or any combination thereof, in any amount, and with any combination of ranges and instants. an instant is just a range where start == end (length 0), no separate type. concatenation goes by length, so an instant adds no duration – it's a marker at the current offset, drawn as a diamond instead of a bar.
because an annotation is just a mark-group, and a timeline is just a mark group, any annotation can be pushed onto the timeline-stack, replacing the main timeline. the clips within the marks in the mark group are laid end to end to form one continuous duration, a new timeline. the cool thing here is that you can now annotate within the context of this annotation. so if we are inside annotation A, annotation A is the :parent of our new Annotation B. if annotation B uses absolute timestamp marks, they are relative to the annotation A timeline, not the main timeline. and if annotation B uses clip based timestamps, they can only reference the frames of the underlying clip which are within range of the annotation (annotation A, the parent, may have start half way through the clip at the beginning, and end half way through the clip at the end). and then you can push annotation B onto the timeline stack, annotate within that, and on and on.
something about mark groups to note is that marks need not be defined in order. if the main timeline has clip A, B, and C laid end to end, an annotation X can have marks /noe/tl/src/commit/7ac8f27b570b9e138076149fd616297779782a2b/tl/clipC%5B0%5D,%20clipC%5B-1, [clipA[0], clipA[-1]], [clipB[0], clipB[-1]]. -1 represents last available frame of the clip (note here that available frame may differ from absolute last frame of the underlying clip, because the annotation context we're in could cut off half the clip, for example). note this clamping only bites for raw clip refs across a trim – if you reference the subclip (the parent's mark) instead, the trim is baked into the mark's range, so subclip[-1] = the subclip's own end, no clamp (more below). in this example, we have totally rearranged the clips into a timeline B C A, end to end. and you can also imagine we can cut clips in half, interleave them, repeat them and so on.
hmm here's a struggle though. let's say i have interleaved half of A with half of B in an annotation which i pushed onto the timeline stack:
marks: /noe/tl/src/commit/7ac8f27b570b9e138076149fd616297779782a2b/tl/clipB%5B20%5D,%20clipB%5B40, [clipA[0], clipA[20]], [clipB[0], clipB[20]] [clipA[20], clipA[40]]]
we want to be able to mark these sub clips independently for a new annotation Y, right? but these are just 2 root clips that became 4. so how can we define annotation Y with respect to any of these 4 clips? we don't want to use absolute timeline time, but we also don't want to use absolute clip time because our marks can cross between the first sub clip (b 20 - 40) and the second (a 0 to a 20), so clip time means nothing here. we also need to remember that the parent timeline can clear out its marks at will. so we can definitely orphan annotations - that's ok, that's a UI concern we can display warnings for and just grey out basically (drop to bottom of annotation list, for example, with a warn emoji and let the user edit to specify its marks again, even warn on which marks are broken, and if at least one mark is still valid still display it there). so it rly seems like an annotation SHOULD create synthetic subclips which point at the raw clips so that we can use the synthetic subclips as our targets. but they need to be stable identities, serializable/deserializable.
resolved: every mark gets a stable id (a uuid) at creation, and a clip/subclip-ref mark stays within ONE clip – so the addressable subclip is just a single-clip mark, addressed by mark-id alone. a selection that crosses clip boundaries is stored as a RUN of per-clip marks (the lane sticks the contiguous run into one bar; the editor splits/merges at boundaries on edit). a mark is already a recursive structure pointing at a raw clip, the id just makes it addressable, so no separate entities, no recreation lifecycle. Y references A's subclips by mark-id: {:ref <mark-id> :at n}.
- edit a mark's range -> same id -> Y follows it
- reorder marks -> ids travel with them -> Y follows the content, not the slot
- delete a mark -> id gone -> Y dangles -> orphan (grey out, warn per broken mark, keep if at least one still resolves)
- add a mark -> new id
so "keep identity unless the whole thing is different" isn't an algorithm we run, it's a ui affordance: editing a row in place keeps the id, delete-row + add-row makes a new id. like keyed list editing / db rows with primary keys – you carry stable keys, you never diff structure to guess identity.
and the subclip bakes the trim into its own definition, so it's a clean map to source: subclip[n] = src-start + n, subclip[-1] = src-end, no clamp, no context lookup. that makes resolving a ref context-free: (mark-id, scene) -> source range, no timeline-stack needed. the stack only decides which context's local timeline you're looking at, not how a ref resolves. only absolute (bare number) points are context-dependent – they're local frames of the context the mark lives in.
so the whole mark grammar is two point kinds:
- ref point {:ref <mark-id> :at n} -> frame n of that mark's resolved range (context-free, n negative = from the end)
- absolute point <number> -> a local frame of the context the mark lives in (resolved through that context's spans)
mix them in a single range, instant = start==end, track optional on the mark. a raw clip is just a mark whose range is full source + a track. a clip/subclip-ref mark keeps both endpoints on the SAME target, so it resolves to exactly one source segment; only an absolute mark may span several.
decided: a clip/subclip-ref mark never crosses a clip boundary, so the "range whose endpoints are in different subclips" case just can't happen. a selection across clips is a run of single-clip marks instead, one per clip, laid end to end (genuinely contiguous – the lane only draws them as one bar). this keeps resolve trivial (one source segment per ref mark), makes every piece addressable by mark-id alone, and makes a reorder follow each piece independently instead of swelling. the merge/unmerge is localized: dragging a boundary WITHIN a clip edits that end mark in place (id stable); dragging ACROSS a clip boundary adds/removes a whole clip from the selection (creates/destroys that end mark); the fully-contained middle clips never churn. the same rule applies at any depth – "clip boundary" means a boundary in the fully-resolved footage, so it works the same whether you're at root crossing raw clips or nested crossing a parent mark's segments. the one multi-segment mark left is an absolute one (bare-number local range): it's arrangement-relative, you build it by dragging the timeline rather than by referencing, and it resolves via slice.
this brings us to playing. since right now there is only one source video file, we need to be able to seek to arbitrary frames. each context (mark-group) keeps its own local playhead, used when it's the top of the timeline stack. when we hit play, in the example above of annotation X, we find the clip under the playhead, compute the source frame, seek there, and start playing. one correction though: the local playhead has to be the master clock, not the video. you can't derive local position from currentTime – once an annotation repeats or reorders clips, one source frame maps to several local frames, it's not invertible. so the local playhead advances on its own (wall-clock x fps while playing), and every frame we compute expected = group->media(local) and seek the video there only if round(currentTime*fps) != expected. within a clip, expected tracks the video's natural playback so no seek fires; at a mark boundary it jumps once and we seek. and right – no recursion at play time: we resolve the current context once into flat ordered spans, and group->media is just the flat lookup the renderer already does.
update!
- ok so the idea is this. you hit the new annotation button. it does not auto-select a mark for you. you can either click the clip, click the frame button, or drag a range. after you select a range, you are automatically in drawing mode. your drawings are connected to the mark, not the annotation. a mark should only ever appear in one annotation. instead of creating a new annotation mark group by default, this mode also allows you to either create new or associate with an existing annotation. associate with existing gives you our dropdown with only other annotations available. when you pick one, you effectively go into "edit" mode on that annotation with the new marks suddenly added. so this is basically our "transclusion": we can have annotations with marks embedded in other timelines. this is great for if we have subdivided our analysis into "chapters" but want to annotate shared concepts across them while keeping the main annotation pane clear. it's organized. so the big thing is that you don't create the annotation, you create the mark(s) first, then either create or assoc the annotation.
- another crucial thing: if we click and drag and it spans multiple clips, the range we have in our create/add annotation UI in the annotation pane should only show the start and end points relative to the clips at start and end. so the way this will work is we will create a mark group that's not an annotation as our "proxy" marks so that they don't appear in the ui that spans the full range and contains the full sequential clips, and then the mark group on the annotation that contains that mark group will just use that mark group as start and end as if it had been clicked. so, for example: there's clips A, B, C and D contiguous. user drags region from clip A to clip D. in the UI, we should see that our mark starts at clip A frame 0 and ends at clip D last frame, so we need to "pass through" the synthetic unnamed non-annotation mark group to the underlying clips, the synthetic unnamed non-annotation mark group is a proxy.. so that proxy mark group has marks that go from clip A start-clip A end, clip B start - clip B end, clip C start to clip C end, and clip D start to clip D end. makes sense? so what do we do if the user wants to drag adjust endpoint in UI? let's say there's another clip before clip A called clip 0. we move start point BACK to clip 0 frame 50. well, our main mark group just shows the range as we would expect: clip 0 frame 50 TO clip D last frame. but the proxy mark group? it has a new mark range with new mark id at the beginning, but the other mark ids are stable. and the same is true of rolling the end point forward: new mark id, new range, rest are stsable. what if we roll the endpoints inward? same principle but we kill off mark ranges instead of adding new ones. should be clean. so this means we needs we need to change how draft/edit marks look/act in the lane. when in draft/edit mode, clicking on the mark brings up drawing mode for that mark (there can still be a button next to the mark in the edit pane). you can drag the whole mark left to right. and you can also grab handles on the edges of the annotation left to right. and since we consider whatever last created or last touched mark to be the "active" one for associating drawings to, we need to have that visually represented in the timeline, and in the pane where the draft marks or edit annotation is. and these need to share the same look for the mark range,s its only the stuff above it that iwll change. make sense?
- note: when we talk about rolling the whole clip around, we know that the mark ids are going to change if we highlight one clip, unhighlight, then return back. this means that if we annotated a range defined w/r/t that annotation, the underlying gids are broken forever, even if they're rolled back. so that they exist still, right? since the underlying clips will never change, i wonder if we could just give each clip a stable identifier and define our root-most ranges in terms of those stable identifiers? or is that worse? idk