Skip to content

Latest commit

 

History

History
407 lines (334 loc) · 23.8 KB

File metadata and controls

407 lines (334 loc) · 23.8 KB

Algorithm and pitfalls

← Documentation ・ ← SmartCut ・ 日本語

The algorithm

For a kept interval [t_in, t_out):

One kept range across a row of pictures: the head, before the first key frame the copy can start at, and the tail, after the last one it reaches, are re-encoded; everything between them is copied byte for byte
... I ....... I=========================I ....... I ...
      ^t_in   ^k_first                  ^k_term   ^t_out
    |<-head->|<--------- body --------->|<-tail->|
      re-encode        stream copy       re-encode

head and tail begin or end partway through a GOP, so there is no choice but to decode from the previous access point and rebuild them. The body between them is emitted as the input's own bytes. Cut exactly on access points and there is no re-encoding at all.

Rebuilding assumes there is an encoder to rebuild with, and for one codec there is none. VC-1, which most Blu-rays pressed before about 2010 are written in, has no encoder in libavcodec, on a graphics card, or in any free implementation. SmartCut writes those pictures itself — intra pictures only, which is all a head or a tail needs, since nothing outside the fragment may be referenced. What that costs, and how it was measured, is in the Rust core.

The Python reference implementation is split as follows:

File What it does
probe.py Stream parameters, the access point index, leading-picture detection and the reference test
planner.py Turns intervals into a segment list
bitstream.py Annex-B / MPEG-2 access unit splitting
renderer.py Runs ffmpeg and concatenates the results
verify.py Decodes the output and compares it against the source frame by frame

Pitfalls

This section is why "just cut on GOP boundaries and join the pieces" is not enough. Every one of these problems was hit for real while building the prototype, and every one has a test pinning the reproduction.

1. The parameter sets (SPS/PPS) do not match

Making the re-encoded part's SPS match the original stream bit for bit is effectively impossible: a different encoder means different VUI and different VBV. And an MP4 avcC box, like a Matroska CodecPrivate, can only hold one set of parameter sets. Concatenate naively and either the copied part or the re-encoded part gets decoded with the wrong SPS, and the picture falls apart.

The fix has two parts:

  • Write each piece as a raw Annex-B elementary stream and join them by plain byte concatenation. An elementary stream carries SPS/PPS in band at every IDR, so parameter sets that differ from piece to piece can legally coexist.
  • Write the final MP4 with an avc3 / hev1 sample entry. That is the in-band form defined by ISO/IEC 14496-15, and it is not folded into avcC.

MPEG-2 does not have this problem in the first place, because its sequence headers are in band already.

A transport stream's record holds one set, and it is the first in the file. In a cut SmartCut wrote, that is the re-encoded head's, under the same ids as the copied pictures' own and with other contents. A re-encode reading such a file seeks, lands ahead of the entry point it aimed at, and the run-up pictures before the next in-band set were parsed with the head's. They were accepted and left wrong references behind: on x264 material with no IDR, every B picture of the re-encoded stretch came out at 28-35 dB in a re-cut of a cut. So nothing before the read's first key packet goes to the decoder (cut::RunUp, in the re-encode, the conformed reel and the crossing's far side), as the preview has always done. A file whose key packets are not marked is fed from the entry point the read aimed at, and one that says nothing from a few hundred pictures in.

2. The access point index has to scan packets

ffprobe -skip_frame nokey misses access points in open GOPs, because the decoder cannot output an I picture whose references are absent. On the prototype's test material it found only 3 of the 10 access points that were actually there.

Looking at the packet's K flag avoids decoding entirely. It is faster, and it is correct.

3. Leading pictures — the heart of the open-GOP problem

Pictures that come after an I picture in decode order but before it in display order are called leading pictures, and they reference the previous GOP. Start a copy there and they cannot be displayed.

Here the handling diverges completely, depending on whether the leading picture is itself a reference picture:

  • MPEG-2: B pictures are never referenced, so leading pictures can simply be dropped. Even an open GOP works as a copy start point.
  • H.264 (x264's open-gop and equivalents): B pyramids mean a leading picture can be a reference picture. Drop one that a later frame predicts from and that frame breaks, taking the rest of the GOP with it. Being a reference does not by itself mean anything after the entry point uses it, though; see below.
  • HEVC: leading pictures are RADL and RASL, and the standard forbids either from being a reference for the trailing pictures of the same entry point. So they can always be dropped, whatever the encoder did — and RASL pictures must be, since what they reference lies on the far side of the cut. That is bitstream::leading_always_droppable, and it is what lets a copy start at the next entry point rather than at the next one nobody measured. A 4K broadcast is the case that needed it: every entry point of one is a CRA, none is an IDR, and measured picture by picture they all read as un-startable. A minute cut out of one re-encoded 13.3 seconds of its head before this and 0.44 after — 22.8% of the range against 1.4%.
  • VC-1: B pictures are never referenced either, so it behaves like MPEG-2 — but its picture header cannot be read on its own. Even whether a picture states its type in three bits or in one is settled in the sequence header, which a transport stream restates in front of every entry point and libavformat hands over as extradata. If no sequence header ever turns up, SmartCut answers conservatively and treats every picture as a reference. The only cost of that is losing the chance to start a copy at an open GOP.

So SmartCut reads nal_ref_idc (H.264), the NAL type (HEVC) or picture_coding_type (MPEG-2) out of the bitstream to tell whether a leading picture is a reference at all. Where none is, the point is a copy start point. For H.264 a leading reference picture no longer settles it. leadrefs answers it the way a decoder would have to: it follows the short-term reference pictures from the entry point on (sliding window, frame_num gaps, memory_management_control_operation 1) and builds the reference lists of every slice of every picture that is not a leading one, as 8.2.4 of the standard does — the initial order, the field alternation, the modifications, the active length.

Nothing naming a leading picture is not enough, because the cut is not the recording. Dropping a leading reference picture leaves a gap in frame_num, and the cut's decoder infers a frame to fill it. That frame takes the dropped picture's slot in the sliding window and in a P picture's list, but its order count is the decoder's choice and lands above the I, so it enters a B picture's default lists where the dropped picture never stood. So the leading pictures are dropped only where all of these hold:

  • No list of a kept picture names a leading picture while it is still held.
  • While a leading reference picture, or a frame inferred after one, is held, every B slice the copy keeps names every active entry of both its lists through ref_pic_list_modification. A list left at its default, even partly, makes the point needed.
  • No leading reference picture carries a memory_management_control_operation of its own (adaptive_ref_pic_marking_mode_flag 0).
  • No leading reference field is held without its other field.
  • The key picture itself is a reference picture.

Every doubt is answered "needed", which is the flag's old answer: long-term references, any other memory management operation, pic_order_cnt_type 1, a parameter set it has not been shown, a slice it cannot read, or a point still undecided after a thousand pictures. Only H.264 whose parameter sets travel in band is followed. A transport stream or .m2ts always qualifies, and so does an MP4 or Matroska file remuxed from one, which keeps its SPS and PPS in the key packets. An MP4 or Matroska file whose parameter sets live only in avcC keeps the flag's answer. The index stores the answer per entry point (seek_index::VERSION 15).

The B-list condition was found the hard way. A first version without it marked entry points droppable on a clip shaped like a pressed Blu-ray's (23.976p H.264: an I, a leading B that is a reference, a leading B that is not). The trailing B named the I explicitly in list 0 but left list 1 at its default length, and the cut failed --verify: two frames off, 16.2 dB. The condition rejects it.

The prototype once measured the opposite of all this: dropping the leading pictures of x264 open-gop material made all 60 frames of the copy's first GOP mismatch, and keeping them made them match. That has not been reproduced. Twenty x264 transport stream variants were tried (b-pyramid normal and strict, 16 references, 8 and 16 B frames, MBAFF top and bottom field first, bluray-compat with 4 slices, weightp 2 and others), each with the leading pictures cut out and decoded from the key, and none of them came out wrong. The rule as it then stood judged all twenty droppable.

The rule now judges them needed, because every leading reference B that x264 writes carries operation 1 of its own. An earlier version let operation 1 through where it unmarked only pictures from before the entry point: the cut still holds those, and they are older than everything after the I. Holding them is the trouble. By order count such a picture stands below everything after the I, so in a trailing B's list 0 it comes in front of every picture shown after that B, and it can stop list 1 from being swapped (8.2.4.2.3). An x264 stream whose headers were edited into a conformant shape, with a trailing B reference that unmarks the I, showed it: the B after that one, on its default lists, decoded differently with and without the leading pictures. So x264 open-GOP material with a B pyramid starts no copy at an open GOP, as when the reference flag alone decided. The recorder's BD-RE below is unaffected, since its leading B frames carry no operation.

The material that needed it is a recorder's field-coded BD-RE. Each GOP has two leading B frames in front of its I, and the first is a reference for the second and for nothing else: the P pictures name the I's fields through an explicit reordering that steps over the pair, and the trailing B pictures' lists are too short to reach back past the I. Judged by the flag, no entry point of such a recording could start a copy, so each range's head was re-encoded up to the first point nobody had read. Titles of that disc went from 90–111 seconds of re-encoding to 5.5–10.

As a copy end point an open GOP is fine either way — the display range simply ends at lead_start.

Removing leading pictures cannot be expressed in the container (it would need an edit list), so the Annex-B access unit boundaries are parsed by hand and the pictures cut out there (bitstream.py). Note that H.264's "first slice = start of picture" rule does not carry over to MPEG-2, where a picture start code is followed by several headers before the slices, so the access unit split differs between them.

4. Specify intervals in frames and packets, not seconds

Pass -t a duration in seconds and it goes wrong twice over:

  • Under -c copy, -t is evaluated against the DTS. DTS runs ahead of presentation time by the reorder depth, so extra I/P pictures from the next GOP sneak in — 180 frames came out as 182.
  • With fractional frame rates (30000/1001), rounding shifts the result by ±1 frame.

The only robust approach was to count the display frames and the packets to copy as integers and pass those (-frames:v N). The packet count follows exactly from the difference of decode-order indices in the access point index.

5. Container start_time (MPEG-TS)

TS timestamps do not start at 0 — the test material starts at 1.423 s. -ss is relative to the start of the file, but ffmpeg re-bases the output by the -ss value alone, so start_time survives as a residual offset in the output timeline.

A seconds-based -t runs straight into that. Access point times are therefore normalised by subtracting start_time, and interval lengths are passed as frame counts, which avoids both problems.

6. Re-encoded regions must be decoded from earlier

With open-GOP sources, seeking straight to the target position with -ss makes ffmpeg discard the entire GOP it could not decode, so the output starts up to one GOP late (0.2 s, measured). Decoding starts from an access point a few GOPs earlier, and the front is trimmed with an output-side -ss.

7. Audio has no GOP structure

Audio is cut per kept interval, not per video segment.

  • --audio-mode copy uses the source frames as they are. Interval boundaries snap to the nearest audio frame, which is up to about 24 ms for AAC.
  • --audio-mode reencode makes one pass through the atrim and concat filters. Sample accurate, but it re-encodes everything.

The Rust core adds a third mode, smart, which is its default: the frames a boundary falls inside are re-encoded and the rest are copied. It also takes a channel count (--audio-channels), which is what folds a 5.1 recording to stereo; that has no copy path at all, so it is a whole-track re-encode whatever the mode says.

Both are covered in detail in audio.

Handling AAC encoder delay and priming strictly via edit lists has not been implemented.

8. The verification reference decodes from the start of the file

The reference side of --verify must not be built with -ss. On open GOPs the reference itself shifts, for the same reason as #6, and a correct cut gets reported as wrong. Decode from the beginning and slice by frame number.

9. Never plan a re-encode window with no picture in it

The tail runs from where the copy's coverage ends to the end of the range, t_out. Those two are usually a few frames apart -- but where t_out lands just after an access point (60.060 s asked for on a 29.97 fps stream whose picture sits at 60.05996 s), the gap is thinner than the spacing between two pictures.

Such a window has no picture to decode. The cutter refuses a window it can decode nothing out of -- otherwise a re-encode that produced no frames would pass in silence -- so the run stops, not just that clip.

A tail is therefore written only when it is at least half a frame wide, for the same reason the head has to be a whole one; the Python reference draws the line in the same place. As a net under both, a re-encode segment whose frame count works out to zero is not emitted at all.

10. The picture order counts either side of a splice are not one another's

A decoder hands its pictures back in picture-order-count order, and the counts on the two sides of a seam were written by different encoders. The re-encoded head opens a coded video sequence of its own and counts from nought; the copied body spliced on after it carries the counts the recording gave it.

Where the copied segment begins on an IDR that settles itself: the sequence restarts and the decoder empties what it was holding. But the discs one authoring tool writes, and a recorder's own, put the other kind of entry point there — an I picture with a recovery point, which restarts nothing. A slice states only the low bits of its count, pic_order_cnt_lsb, and the decoder derives the rest from the reference picture before it, so the copied I is read against the last low bits the head wrote. Where it derives lower than pictures already shown, libavcodec takes the copy's first pictures for ones it has put out and drops them: 7 of 714 on an authored disc, 15 of 1379 on a recorder's disc coded as field pairs, and a run of twelve cuts that decoded 698 of its 718. The copy itself was never wrong — the same pictures decoded from the recording from that same I come out bit for bit.

So the head is renumbered, and only where it has to be. cut::poc_seam reads the head's slices and the copy's first one and derives what the copy's I would be shown at. Where that is not after the head, every slice of the head has its pic_order_cnt_lsb moved along by one amount, chosen so that the I lands past the head's last reference picture by 2 × max_num_ref_frames of the copy's sequence. That margin is for the frames libavcodec invents to fill the frame_num gap at the seam: left above the I, they sort into the lists its B pictures are predicted from, and on the recorder's disc two apart still left 2 of 718 copied frames wrong. The differences between the head's own pictures stay what they were, and the copy is not touched.

In order is not enough; clear of the invented frames is the test. A copy that already came after the head used to be let through as it was, and the margin was never checked: a range opening on an entry point whose leading pictures are dropped, after a head of a dozen pictures, had its two B frames predicted from invented frames (--verify: 2 of 1183 copied frames differ, 28.5 dB). So the shift is chosen in this order: clear at the field's own width, clear at a byte wider, and only then merely in order. A copy that is already in order and cannot be made clear is left as it is.

The field is sometimes too narrow to say it. x264 writes as few bits of order as its own pictures need: on the recorder's disc, where the head is MBAFF, four against the recording's eight, which cannot reach more than eight frames ahead. Where only a wider field makes the copy clear, the head's SPS widens log2_max_pic_order_cnt_lsb by 8 and every head slice's field grows by a byte. A whole byte, because an arithmetic-coded slice starts its data on a byte boundary, and moving everything after the field by exactly one byte keeps that boundary without the rest of the header having to be read.

A range's re-encoded tail followed by the next range's copy of the same recording is the same seam and is handled the same way. The decision needs the copy's first picture, so the writer holds every write to the muxer — sound included — from the start of the head until that picture arrives, and lets them go in the order they were made; a cut that needs no renumbering is byte-identical to one made without any of this. An earlier attempt (2026-09-09) rewrote the head's counts unconditionally, made things worse, and was dropped; this one acts only where the reversal, or an invented frame above the copy's I, is derived, and was measured never to lower a decoded count. Not covered: the seam between two reels of a join with no transition, and HEVC, whose CRA entry points have not been tried.

--clean-joins remains as an option. It spends a little more re-encoding to reach an IDR instead, which restarts the sequence. What it costs is exactness: the stretch between the two entry points stops being copied and becomes a re-encode, measured at 51 dB against a decode of the same pictures with their own references. That is a good re-encode and it is still not the recording, so this is off unless it is asked for.

How far it is worth reaching is the other half of that option. examples/idrdiag.rs walks a recording and prints the wait for a clean entry point from each of its own, and the four H.264 discs to hand disagree completely: two wait a second or three at the median, and one holds a single IDR across a twenty-six minute title, where the median wait is thirteen minutes. So the reach is bounded at two seconds, the plan is left exactly as it was wherever nothing clean is within that, and material that reorders nothing — MPEG-2 and VC-1, which state each picture's display order within its own group — never moves at all.

Reading this needs the recording in hand, and on a disc the index was never walked: it came off the disc's own table. So planning gained an entry point that has the recording, plan_on, while plan stays the arithmetic. It costs a handful of seeks rather than a pass — the byte each entry point begins at is already known.

11. A recording may hand over two pictures per frame

The output timeline counts in fields, because pulldown shows some pictures for three of them and a frame grid cannot say so. That answers how long a picture stands. It does not answer how many pictures a frame arrives in, and the answer is not always one: H.264 lets a frame be coded as a pair of field pictures — PAFF — and a recorder writing its own Blu-ray discs does exactly that, 1440x1080 at 29.97 with every frame split in two.

Broadcast 1080i is interlaced content coded as whole frames, so nothing in the broadcast corpus hits this. On a recorder's disc it breaks the cut in two places at once:

  • Placement. Each field is half a frame and takes one field of the timeline, not two. Given two, the segment's span is twice what the pictures occupy and the duration written onto each packet is twice what it lasts.
  • Lookahead. DTS is derived by holding pictures back until the display order is known, and how many to hold comes from the recording's reorder depth, which counts frames. Held to one picture, the writer settles a P frame's decode time before the B frames that display ahead of it have arrived. Every one of those is then left with nowhere to go and dropped: on a forty-second cut, 642 of 1842 pictures, and the copy decoded as blocky mush from its first picture on.

So the queue is measured in halves of a frame rather than in pictures. A picture standing for a single field contributes one half, anything else two — including a picture shown for three fields, which is still one frame and must not move the boundary. Frame-coded material is then held exactly as it was before, and a recording that is frame-coded but for a handful of field pairs gets the extra room only where those pairs are.

The question is asked of the packet, not of the pictures in it. libavcodec's MPEG-2 parser joins a complementary field pair into one packet, and one broadcast in the sample does that seven times in half a minute: two field pictures, one frame, two fields of the timeline. Its H.264 parser hands each field over on its own. So what is counted is the pictures a packet opens — first_mb_in_slice of nought for H.264, one picture coding extension for MPEG-2 — and a packet is half a frame only where it opens one picture and that picture is a field.

Reading the flag needs the sequence header first: field_pic_flag sits behind frame_num, whose width only the SPS carries. That is read once per recording, the way a VC-1 stream's headers are, and a sequence with frame_mbs_only_flag set needs no picture asked at all. The header alone does not settle it, though: what this program writes back over the same recording is MB-AFF, whose header also allows fields while every picture in it is a whole frame.

The second field of a pair is placed after the first rather than by its own timestamp. The two are always next to each other in decode order, so the first is always the one just seen; a recording that gives a pair one timestamp between them — or two close enough to round together — would otherwise put both fields in the same place, and one of them would have nowhere to go.