Skip to contents

Normalize the audio loudness of many input files (EBU R128) from a single jobs tibble. This is the batch (table-driven) form of normalize_audio(), for when you have more than one file to normalize. Each row is one input, and the only required column names its source. The function is a thin wrapper over ffm_batch. It gives one reproducible compiled command per input. Each row uses the same loudnorm pipeline, and the same check of each value, as normalize_audio(). Set two_pass = TRUE for accurate measured/linear normalization across the whole table (see two_pass). The glossary in vignette("tidymedia") explains media terms such as LUFS, encoder and sample rate.

Usage

normalize_audio_batch(
  jobs,
  target_loudness = -23,
  true_peak = -1,
  loudness_range = 7,
  channels = NULL,
  sample_rate = NULL,
  audio_codec = NULL,
  two_pass = FALSE,
  audio_stream = NULL,
  run = TRUE,
  parallel = FALSE,
  ...
)

Arguments

jobs

A data frame with one row per input and (at least) an input column (source path). An optional output column names the destination. Without it, the function derives one per row. It appends _normalized to the basename of each input and keeps the extension of the input (e.g. clip.mkv becomes clip_normalized.mkv). The derived name keeps a video extension while the file itself holds audio only. So name an output column yourself when that matters. The function refuses two rows that name the same output path before any row runs. With two_pass = TRUE, that is before the analysis pass. The refusal covers a path repeated in the output column, or a repeated input when there is no output column. Five loudness arguments can also appear as a column that overrides the argument per row. They are target_loudness, true_peak, loudness_range, channels and sample_rate. Rows that omit the column fall back to the value of the argument. An optional audio_codec column (character) names the output audio encoder of each row. There, NA means "leave the encoder unset". Rows that omit it fall back to the audio_codec argument. An optional numeric audio_stream column likewise overrides the audio_stream argument per row. There, NA normalizes the first audio track of that row. Any other columns are ignored.

target_loudness, true_peak, loudness_range

The EBU R128 loudness targets applied to every row, unless jobs carries a column of the same name (see jobs). Defaults follow EBU Recommendation R 128 (2014): target_loudness = -23 LUFS, true_peak = -1 dBTP, loudness_range = 7 LU.

channels

The output channel count applied to every row, unless jobs carries a channels column, e.g. 1 to downmix to mono. NULL (default) keeps each source's channel layout.

sample_rate

The output sample rate in Hz applied to every row, unless jobs carries a sample_rate column. NULL (default) lets loudnorm choose. It resamples, up to 192 kHz and capped by the encoder, and does not keep the source rate. Set this argument to fix the output rate.

audio_codec

The output audio encoder applied to every row, unless jobs carries an audio_codec column, e.g. "aac". NULL (default) sets no -codec:a, which leaves the default encoder of the output container in place. "copy" is an error. Loudness normalization filters the audio, so it must be re-encoded. See normalize_audio.

two_pass

A logical that selects the normalization mode for every row. It applies to the whole table and is not a per-row column. FALSE (default) keeps the single-pass loudnorm pipeline. TRUE runs accurate two-pass (measured/linear) normalization in two phases. An analysis pass first measures the loudness of every input. It honors parallel and the targets of each row. A correction pass then feeds those measurements back with linear=true, so each output hits its EBU R128 target precisely. This is the table-wide form of two_pass in normalize_audio. The result shows the five measured values as columns measured_I, measured_TP, measured_LRA, measured_thresh and offset. Two-pass must measure each input. So it always runs the analysis pass through FFmpeg, even when run = FALSE. It needs the binary and readable inputs. If the analysis of any row fails or gives no measurement that can be parsed, the call aborts and names those rows. It aborts before it builds any correction command. That abort has class tidymedia_loudnorm_no_measurement, the same class that normalize_audio raises for this event. It carries the same row numbers on tm_rows, alongside tm_row_status. That field has the FFmpeg exit status of each row, or NA where the row exited zero but printed nothing that can be parsed. It carries no single exit status on tm_status, and it does not have class tidymedia_ffmpeg_exit. The reason is that it also fires for rows that exited zero. A batch can mix causes, so there is no one number to report. The one-file form carries both only where FFmpeg exited non-zero. Where FFmpeg exited zero and printed nothing that can be parsed, the one-file abort carries the shared class alone, with no tm_status either. Silent rows are the exception. A silent input (analysis loudness -inf) cannot be normalized to a target, but one silent row does not abort the batch. The function normalizes the rows that are not silent. It marks the silent rows in a logical silent column, with success = FALSE and no output written, and a warning names them. This is where the batch form and the one-file form differ. normalize_audio aborts on a silent input, because one silent input is the whole call. Here, the other rows still have work to do. The single-pass default touches no binary under run = FALSE.

audio_stream

The audio track to normalize, as a number that counts from 0 among the audio tracks of each row's input. 0 is the first audio track and 1 is the second. Other streams in the file, such as video, do not count. NULL (default) normalizes the first audio track. Without an audio_stream column, the argument applies to every row. An NA cell in that column means NULL for that row. It does not fall back to the argument. The first-track family reads NULL as the first audio track only: extract_audio, convert_audio and normalize_audio, and their _batch forms. The every-track family reads it as every audio track: separate_audio_video, standardize_video, anonymize_video, crop_video, segment_video and format_for_web, and their _batch forms. This function reads NULL as the first track only. The two-pass analysis measures each audio track, but the correction uses one set of measurements. Normalizing several tracks at once would apply one track's measurements to all of them. Under two_pass = TRUE, the analysis pass measures this same track. Only the named track reaches the output, and no video does, whatever the container. So an output name with a video extension gives a video file that holds only audio. An input with no audio at all is an FFmpeg error. A track the input does not have gives an FFmpeg error, not an R one. See audio_stream for how this differs from audio_input, the input index on compare_videos and picture_in_picture. (default = NULL)

run

A logical: run each input's command through FFmpeg (TRUE, default) or only compile them for inspection (FALSE). Under two_pass = TRUE this gates only the correction pass; the analysis pass runs regardless (see two_pass).

parallel

A logical passed to ffm_batch: normalize in parallel with furrr (TRUE) or sequentially (FALSE, default). Parallelism follows the active future plan. TRUE under the default sequential plan runs one input at a time and warns. Set a plan first, e.g. future::plan(future::multisession).

...

Additional arguments forwarded to ffm_batch, such as verify, manifest, checksums, and progress.

Value

The tibble returned by ffm_batch: jobs with an added command column. When output was derived, it also has the resolved output column. When run = TRUE, it has a success column, plus any columns the forwarded arguments add, such as verified. Under two_pass = TRUE, the result also carries the five measured columns (measured_I etc.) and a logical silent column. The command column then holds the linear correction commands. It is NA for silent rows, which carry NA measurements and are not normalized. The columns of the two-pass result do not depend on how many rows are silent. The verified column (under verify) and the provenance manifest (under manifest, read with ffm_manifest) are present whenever requested. That holds even when every row is silent. Silent rows carry NA for those outputs.

Details

The function warns once for the whole batch when a row names no audio_stream and its input carries tracks that the output will not. The warning names every affected row. Naming a track silences it. Use the audio_stream argument, or an audio_stream cell on every row. suppressWarnings(classes = "tidymedia_dropped_audio") silences it too. The check costs one FFprobe call per distinct input it has to probe. A repeated input is probed once, and a row that names a track is not probed at all. The warning is given when FFprobe is available and the input can be probed. Otherwise the check is skipped silently. Those probes run one at a time, before any row starts, so parallel does not reach them. A sweep long enough to look like a hang reports its progress. The check never runs under run = FALSE and never changes any compiled command. It is skipped entirely when every row names a track. Under two_pass = TRUE, the warning comes before the analysis pass. So it arrives while adding audio_stream can still save that pass.

To switch the check off and skip the whole sweep, use options(tidymedia.check_tracks = FALSE) for the session. Use withr::local_options(tidymedia.check_tracks = FALSE) for the rest of one function.

References

EBU Recommendation R 128 (2014), Loudness normalisation and permitted maximum level of audio signals; ITU-R BS.1770-4.

Examples

video <- system.file("extdata", "sample.mp4", package = "tidymedia")
jobs <- tibble::tibble(
  input           = c(video, video),
  output          = c(tempfile(fileext = ".m4a"), tempfile(fileext = ".m4a")),
  target_loudness = c(-23, -16)
)
# run = FALSE compiles one command per input without calling FFmpeg
normalize_audio_batch(jobs, run = FALSE)
#> # A tibble: 2 × 4
#>   input                                           output target_loudness command
#>   <chr>                                           <chr>            <dbl> <chr>  
#> 1 /home/runner/work/_temp/Library/tidymedia/extd… /tmp/…             -23 "-y -i…
#> 2 /home/runner/work/_temp/Library/tidymedia/extd… /tmp/…             -16 "-y -i…
# \donttest{
# Accurate two-pass (measured/linear) normalization across the whole table.
# This one runs FFmpeg to measure each input, so it needs the binary.
if (nzchar(Sys.which("ffmpeg"))) {
  normalize_audio_batch(jobs, two_pass = TRUE)
}
#> # A tibble: 2 × 11
#>   input               output target_loudness measured_I measured_TP measured_LRA
#>   <chr>               <chr>            <dbl>      <dbl>       <dbl>        <dbl>
#> 1 /home/runner/work/… /tmp/…             -23      -21.8       -17.7            0
#> 2 /home/runner/work/… /tmp/…             -16      -21.8       -17.7            0
#> # ℹ 5 more variables: measured_thresh <dbl>, offset <dbl>, silent <lgl>,
#> #   command <chr>, success <lgl>
# }