Skip to contents

Cut segments across many input files from a single jobs tibble. This is the batch (table-driven) form of segment_video(), for when your segments span more than one input. Each row is one segment; the four required columns name its source, destination, and cut points. This is a thin wrapper over ffm_batch: one reproducible compiled command per segment. The glossary in vignette("tidymedia") explains media terms such as codec, keyframe and stream copy.

Usage

segment_video_batch(
  jobs,
  reencode = TRUE,
  video_codec = NULL,
  audio_codec = "copy",
  hardware = c("none", "nvenc", "videotoolbox"),
  fallback = FALSE,
  quality = NULL,
  audio_stream = NULL,
  run = TRUE,
  parallel = FALSE,
  ...
)

Arguments

jobs

A data frame with one row per segment and (at least) the columns input (source path), start and end (cut points). Each cut-point column is a numeric column of seconds or a character column with time-duration syntax. Two optional columns are recognized: output (destination path) and reencode (a logical; see the reencode argument). If output is absent, the function derives one per row by appending _<n>.<ext> to each input's basename. The segment number restarts at 1 for each input file (the same rule as segment_video). Two rows given the same output path are refused before any row runs. A video_codec or audio_codec column overrides that argument per row, with NA meaning "leave the codec unset" (the column's way of writing the argument's NULL). An audio_stream column likewise overrides that argument per row, with NA meaning "keep every audio track" (the column's way of writing that argument's NULL). A numeric quality column overrides the quality argument per row (see quality). Any other columns are ignored.

reencode

A logical passed to ffm_seek: cut each segment frame-accurately by re-encoding (TRUE, default) or with a fast, lossless copy that snaps to keyframes (FALSE). See ffm_seek for the trade-off. Applies to every row, unless jobs carries a reencode column, which overrides this argument on a per-row basis.

video_codec

A string naming the output video codec, applied to every row lacking a video_codec column. NULL (default) leaves it unset, so each segment keeps its container's default encoder. A row can resolve to a codec while cutting by stream copy (reencode = FALSE, as an argument or a column). That row is an error, because no encoder runs on that path.

audio_codec

A string naming the output audio codec, applied to every row lacking an audio_codec column. "copy" (default) stream-copies the audio; name an encoder to re-encode it, or NULL to leave the codec unset. A row can resolve to anything but "copy" while cutting by stream copy (reencode = FALSE, as an argument or a column). That row is an error. So split a jobs table that mixes stream-copy rows with a re-encoding audio_codec into separate calls.

hardware, fallback

The encoder backend and its fallback behavior, applied to the whole batch. They are a property of the machine, not of a row, so neither is read as a jobs column. See segment_video(). Because hardware is batch-wide, a non-"none" value conflicts with a stream-copy row on its own, even one naming no codec. So split a jobs table that mixes reencode = FALSE rows with GPU encoding into separate calls. Resolving a hardware backend asks this FFmpeg build which encoders it has. So the first such call that re-encodes the video runs FFmpeg while the command is built, even under run = FALSE. The answer is remembered for the rest of the R session. See refresh_ffmpeg_capabilities to discard it. This function checks that the encoder is available before any row runs. So an unavailable encoder aborts naming this function, not the internal step that runs the rows. A call can also contradict itself by asking for GPU encoding on a cut that stream-copies. Such a call is refused for the contradiction first, whether or not this machine has the encoder. The stream-copy conflict named under reencode is caught first, so such a call aborts without probing.

quality

A number, or NULL (default), applied to each row unless jobs carries a numeric quality column. In that column, NA leaves that row's encoder default in place, whatever the argument says. The value is the encoder's own rate-control value, passed through unchanged. Each cell is checked against the encoder its own row resolves to. A wrong cell is refused before any row runs, and the error names this function and the row. See segment_video() for the encoders, their flags and ranges, and the values it refuses. A cell on a row that stream-copies (reencode = FALSE) is refused too, because no encoder runs on that row.

audio_stream

The audio track to carry into each output, as a number that counts from 0 among the audio tracks of each row's input. 0 is the first audio track and 1 is the second. Other streams in the file, such as video, do not count. NULL (default) carries every audio track. Without an audio_stream column, the argument applies to every row. An NA cell in that column means NULL for that row. It does not fall back to the argument. The every-track family reads NULL as every audio track: separate_audio_video, standardize_video, anonymize_video, crop_video, segment_video and format_for_web, and their _batch forms. The first-track family reads it as the first audio track only: extract_audio, convert_audio and normalize_audio, and their _batch forms. The function does not carry subtitle or data streams in either case. A track the input does not have gives an FFmpeg error, not an R one. See audio_stream for how this differs from audio_input, the input index on compare_videos and picture_in_picture. (default = NULL)

run

A logical: run each segment's command through FFmpeg (TRUE, default) or only compile them for inspection (FALSE).

parallel

A logical passed to ffm_batch: cut segments in parallel with furrr (TRUE) or sequentially (FALSE, default). Parallelism follows the active future plan; TRUE under the default sequential plan runs one segment at a time and warns. Set a plan first, e.g. future::plan(future::multisession).

...

Additional arguments forwarded to ffm_batch, such as verify, manifest, checksums, and progress.

Value

The tibble returned by ffm_batch: jobs with an added command column. When output was derived, it also has the resolved output column. When run = TRUE, it has a success column, plus any columns the forwarded arguments add, such as verified.

References

https://ffmpeg.org/ffmpeg-utils.html#time-duration-syntax

Examples

video <- system.file("extdata", "sample.mp4", package = "tidymedia")
jobs <- tibble::tibble(
  input  = c(video, video),
  output = c("a.mp4", "b.mp4"),
  start  = c(0, 0.5),
  end    = c(0.5, 1)
)
# run = FALSE compiles one command per segment without calling FFmpeg
segment_video_batch(jobs, run = FALSE)
#> # A tibble: 2 × 5
#>   input                                               output start   end command
#>   <chr>                                               <chr>  <dbl> <dbl> <chr>  
#> 1 /home/runner/work/_temp/Library/tidymedia/extdata/… a.mp4    0     0.5 "-y -i…
#> 2 /home/runner/work/_temp/Library/tidymedia/extdata/… b.mp4    0.5   1   "-y -i…