Skip to contents

Cover fixed rectangular regions of many input videos with opaque filled boxes from a single jobs tibble. This is the batch (table-driven) form of anonymize_video(), for when you have more than one video to redact. Each row is one input with its own regions. The required columns name the source (input) and the boxes to cover (regions). This is a thin wrapper over ffm_batch. It compiles one reproducible command per input. It shares the same box-fill pipeline (and per-region validation) as anonymize_video(). The glossary in vignette("tidymedia") explains media terms such as codec, pixel format and stream copy.

Usage

anonymize_video_batch(
  jobs,
  color = "black",
  video_codec = "libx264",
  audio_codec = "copy",
  pixel_format = "yuv420p",
  hardware = c("none", "nvenc", "videotoolbox"),
  fallback = FALSE,
  quality = NULL,
  audio_stream = NULL,
  run = TRUE,
  parallel = FALSE,
  ...
)

Arguments

jobs

A data frame with one row per input and (at least) an input column (source path) and a regions list-column. Each regions cell is itself a data frame of boxes for that input. It has the same shape that anonymize_video takes: x, y, width, height and an optional per-box color. An optional output column names the destination. When it is absent, the function derives one per row. It appends _anonymized to each input's basename and keeps the input's extension (e.g. clip.mkv becomes clip_anonymized.mkv). Two rows naming the same output path are refused before any row runs. That is a path repeated in the output column, or a repeated input when there is no output column. Four encoding arguments may also appear as a column: color, video_codec, audio_codec and pixel_format. Such a column overrides the corresponding argument on a per-row basis. Rows (or arguments) that omit the column fall back to the argument's value. In either codec column, NA leaves that row's codec unset. This is the column form of video_codec = NULL / audio_codec = NULL. In a color or pixel_format column NA is an error, because those have no unset state. An audio_stream column overrides the audio_stream argument per row, where NA keeps that row on every audio track. A numeric quality column overrides the quality argument per row (see quality). Any other columns are ignored.

color

A string naming the default fill color (FFmpeg color syntax) applied to every row. A color column in jobs, or a box that supplies its own color, overrides it. (default = "black")

video_codec

A string naming the output video codec applied to every row, unless jobs carries a video_codec column. In that column, NA in a cell leaves that row's codec unset. The default is "libx264". NULL emits no -codec:v and lets the output container's default encoder decide. For a .webm output, pass audio_codec = NULL too, because the default "copy" would otherwise carry a codec WebM cannot hold.

audio_codec

A string naming the output audio codec applied to every row, unless jobs carries an audio_codec column. In that column, NA in a cell leaves that row's codec unset. "copy" (default) stream-copies the audio through untouched. Name an encoder (e.g. "aac") when the source audio cannot be copied into the output container.

pixel_format

A string naming the output pixel format applied to every row, unless jobs carries a pixel_format column. (default = "yuv420p")

hardware

The encoder backend applied to every row. "none" (default) uses the software video_codec. "nvenc" gives NVIDIA GPU encoding (H.264, HEVC and AV1). "videotoolbox" gives Apple GPU encoding (H.264 and HEVC). Batch-wide (a machine property), not a per-row column; a hardware column in jobs is ignored. See has_hardware_encoder. Resolving a hardware backend asks this FFmpeg build which encoders it has. So the first such call that re-encodes the video runs FFmpeg while the command is built, even under run = FALSE. The answer is remembered for the rest of the R session. See refresh_ffmpeg_capabilities to discard it. This function checks that the encoder is available before any row runs. So an unavailable encoder aborts naming this function, not the internal step that runs the rows. A call can also be wrong about a per-row value, for example a regions table that is missing a required column. The function refuses that call for the value first, whether or not this machine has the encoder.

fallback

A logical applied to every row. When a hardware other than "none" is requested but its encoder is unavailable, TRUE re-encodes with the software video_codec and a message. FALSE (default) aborts instead. It is batch-wide, not a per-row column. A video_codec in a family that the backend has no encoder for is a wrong argument, not an absent encoder. So it aborts whatever fallback says.

quality

A number, or NULL (default), applied to each row unless jobs carries a numeric quality column. In that column, NA leaves that row's encoder default in place, whatever the argument says. The value is the encoder's own rate-control value, passed through unchanged. Each cell is checked against the encoder its own row resolves to. A wrong cell is refused before any row runs, and the error names this function and the row. See anonymize_video() for the encoders, their flags and ranges, and the values it refuses.

audio_stream

The audio track to carry into each output, as a number that counts from 0 among the audio tracks of each row's input. 0 is the first audio track and 1 is the second. Other streams in the file, such as video, do not count. NULL (default) carries every audio track. Without an audio_stream column, the argument applies to every row. An NA cell in that column means NULL for that row. It does not fall back to the argument. The every-track family reads NULL as every audio track: separate_audio_video, standardize_video, anonymize_video, crop_video, segment_video and format_for_web, and their _batch forms. The first-track family reads it as the first audio track only: extract_audio, convert_audio and normalize_audio, and their _batch forms. The function does not carry subtitle or data streams in either case. A track the input does not have gives an FFmpeg error, not an R one. See audio_stream for how this differs from audio_input, the input index on compare_videos and picture_in_picture. (default = NULL)

run

A logical: run each input's command through FFmpeg (TRUE, default) or only compile them for inspection (FALSE).

parallel

A logical passed to ffm_batch: anonymize in parallel with furrr (TRUE) or sequentially (FALSE, default). Parallelism follows the active future plan; TRUE under the default sequential plan runs one input at a time and warns. Set a plan first, e.g. future::plan(future::multisession).

...

Additional arguments forwarded to ffm_batch, such as verify, manifest, checksums, and progress.

Value

The tibble returned by ffm_batch: jobs with an added command column. When output was derived, it also has the resolved output column. When run = TRUE, it has a success column, plus any columns the forwarded arguments add, such as verified.

Examples

video <- system.file("extdata", "sample.mp4", package = "tidymedia")
jobs <- tibble::tibble(
  input   = c(video, video),
  output  = c("a.mp4", "b.mp4"),
  regions = list(
    data.frame(x = 10, y = 10, width = 120, height = 90),
    data.frame(x = 200, y = 150, width = 80, height = 60)
  )
)
# run = FALSE compiles one command per input without calling FFmpeg
anonymize_video_batch(jobs, run = FALSE)
#> # A tibble: 2 × 4
#>   input                                                   output regions command
#>   <chr>                                                   <chr>  <list>  <chr>  
#> 1 /home/runner/work/_temp/Library/tidymedia/extdata/samp… a.mp4  <df>    "-y -i…
#> 2 /home/runner/work/_temp/Library/tidymedia/extdata/samp… b.mp4  <df>    "-y -i…