A plain-language reference for the terms that recur across these articles. Each entry defines the idea once, and the other articles link here rather than re-explaining it. Terms are listed alphabetically. Nothing here is new: it is the vocabulary the Getting started and Choosing an ICC guides use, gathered in one place.
Absolute agreement
One of the two types of ICC. Absolute agreement asks
whether raters give the same score. Two raters who rank
subjects identically but sit a full point apart do not
agree in this sense. It counts systematic rater differences (the rater
variance) as error. Contrast consistency. In
generalizability theory the absolute-agreement ICC is the
dependability coefficient.
Average-unit ICC: ICC(*,k)
The reliability of the mean of k
raters, rather than of one rater. Averaging cancels part of the error,
so ICC(*,k) is always at least as high as the single-rater
ICC(*,1). The k is the number of raters whose
average you will actually use. On incomplete data it becomes the
effective number of ratings. See single-unit
ICC for its counterpart.
Burch interval
An opt-in closed-form confidence interval for the balanced one-way
design (ci_method = "burch"). It gives REML-based limits
with a kurtosis adjustment (Burch 2011), so its width tracks the tail
weight of the data, where the exact-F interval leans on
normality instead. Tracking tail weight is not the same as widening: on
the two grids this package has measured that vary only the subject
effect it came out the narrower of the two in nearly every cell. How
much narrower is conditional, and not in the direction one might guess.
"burch"’s width margin holds much the same up to a true ICC
of 0.3 rather than shrinking as the true ICC rises (on the larger grid;
the smaller grid’s margin does shrink across its levels). The
"burch" width margin then collapses to near parity at a
true ICC of 0.6, on the one grid reaching that value, where every cell
favouring the exact-F interval sits. It shrinks steadily as the subject
count grows, measured at 5 raters. Burch reports the reverse for
symmetric heavy-tailed data with non-normal errors, and a third grid now
measures that case. What "burch" does against
"searle" depends on what the residual is drawn from. The
three grids now measure that: the two grids that vary only the subject
effect put it narrower nearly everywhere, while the third, which draws
the residual from the same family as the subject effect, puts it wider
at every symmetric heavy-tailed family measured (a median width ratio of
1.2963 at t(5) with 100 subjects) and narrower at every lighter-tailed
one, the normal included. So neither interval is reliably the tighter
one. That adjustment is not a remedy for heavy tails: on strongly skewed
subject effects "burch" under-covers about as badly as the
default. See When the
default under-covers. Deterministic: no resamples, no seed. See
Confidence-interval
methods.
Confidence interval vs. credible interval
A confidence interval (the frequentist engines’
output) is a range built so that, across many hypothetical repetitions
of the study, 95% of such ranges would contain the true ICC. A
credible interval (the Bayesian brms
engine’s output) is a range that holds 95% of the posterior probability.
You can say directly “there is a 95% chance the ICC lies in here, given
the data and the prior.” They answer subtly different questions. See Confidence-interval
methods.
Conflated ICC
The single-level ICC you would get by ignoring a
clustering structure (pupils in classrooms, patients in clinics). This
quantity is ten Hove et al.’s (2022) Equation 14. It folds the
between-cluster and within-cluster variation into one “true score” and
is biased for both the subject-level and cluster-level questions.
icc() can report it (level = "conflated")
purely as a diagnostic contrast, to show the cost of
ignoring the structure. It is never a number to report. See Multilevel
designs.
Connectedness (identification)
A design is connected when the raters and subjects
are linked tightly enough that the model can separate a subject effect
from a rater effect. With enough missing cells a design can split into
disconnected islands, and then the variance components are not
identified, and no method can estimate them.
icc() checks this and aborts loudly rather than return a
number that isn’t estimable.
Consistency
The other type of ICC. Consistency asks whether raters
rank subjects the same way, forgiving a constant offset between
raters. Two raters who agree on the ordering but differ by a fixed point
still count as perfectly consistent. It leaves the rater main effect out
of the error term. Contrast absolute agreement.
Dependability coefficient
The generalizability-theory name for the absolute-agreement ICC, written . Projecting it to a different number of raters (a D-study) is a change of the averaging divisor. See D-studies and within-cell replicates.
D-study (decision study)
A forward-looking projection: given the variance components you already estimated, how reliable would the mean of some other number of raters (or occasions) be? It reuses the existing fit, with no refitting, and answers “how many raters do I need?”. See D-studies and within-cell replicates.
Effective number of ratings: k_eff
On incomplete data, subjects are rated different
numbers of times, so there is no single “k” to average
over. k_eff is the harmonic mean of the
per-subject rating counts, and it is the divisor icc() uses
for ICC(*,k) on ragged data. It is always at or below the
full panel size, and the report names it so the divisor is never a black
box.
Engine
The computational backend icc() uses to estimate the
variance components, chosen with the engine argument:
glmmTMB (the default mixed model),
lme4 (an alternate mixed-model solver),
lavaan (a structural-equation formulation), or
brms (a Bayesian fit). Some engines are just a
different solver for the same estimator. Others compute a genuinely
different, though asymptotically equivalent, estimator. See Estimation engines.
Estimand
The true quantity you are trying to estimate: the target the ICC is aiming at. The word matters because “the ICC” is not one number. Agreement and consistency, single and average, subject level and cluster level are different estimands, and picking the coefficient is really picking which one answers your question. See Choosing an ICC.
Exact-F interval
An opt-in closed-form confidence interval for the balanced one-way
design (ci_method = "searle"): it inverts the exact-F pivot
of the one-way ANOVA (Searle 1971, Ch. 9 Table 9.14; the McGraw &
Wong 1996 Table 7 limits). Exact under normality. It is best-calibrated
when the data are approximately normal, and it is deterministic. Exact
is not the same as shortest: see the Burch interval
above for what this package measured about their relative widths. See Confidence-interval
methods.
FIML
Full-information maximum likelihood: the technique the lavaan (SEM) engine uses to fit incomplete data. Rather than dropping cases with missing cells, it uses every observed value to estimate the model. See Estimation engines.
Finite-population rater variance: θ²_r
Raters may be treated as fixed, meaning the observed raters are the whole population of interest. Then the “rater variance” is the spread of just those raters, computed as a bias-corrected finite-population quantity (McGraw & Wong’s Case 3A) rather than an estimate of a wider rater universe. On balanced data it equals the random-rater variance. Under imbalance it differs. See fixed vs. random raters.
Fixed vs. random raters
The raters argument. Random raters (the
recommended default) treat the raters you used as a sample from a larger
pool, so the reliability generalizes to new raters drawn from
that pool. Fixed raters treat the observed raters as
the entire population of interest, so the reliability speaks only to
these raters. The choice changes the rater term from a
random-sample variance to the finite-population rater
variance.
Harmonic mean
An average that leans toward the smaller values: the reciprocal of the mean of the reciprocals. It is the right average for the effective number of ratings because reliability depends on the rate of information per subject, which the harmonic mean captures.
Indicator-mean estimator
How the lavaan (SEM) engine recovers the rater variance for absolute agreement. In lavaan’s formulation a rater is a single column with no random effect, so its variance is read from the spread of the estimated column (indicator) means (Jorgensen 2021). It is a genuinely different estimator than the mixed model’s random effect, though asymptotically equivalent to it, and can differ modestly on small designs. See Estimation engines.
Modified profile likelihood
An opt-in deterministic confidence interval for the balanced,
complete two-way random absolute-agreement design
(ci_method = "mpl"; Xiao & Liu 2013). It profiles the
likelihood in the ICC with a calibrated small-sample correction, and
returns an interval at the near-zero boundary where the Monte-Carlo
default aborts. Available at conf_level 0.90, 0.95, and
0.99, and deliberately conservative. See Confidence-interval
methods.
Monte-Carlo interval
The default confidence-interval method
(ci_method = "montecarlo"). It draws many parameter vectors
from the fitted model’s estimated covariance, on a scale that respects
the zero-variance boundary, recomputes the ICC for each draw, and takes
the 2.5% and 97.5% quantiles. Fast and boundary-aware. It does assume
the fitted parameters are approximately normally distributed around the
truth, and it under-covers when the subject effects are strongly skewed
or heavy-tailed. See When the
default under-covers. See also Confidence-interval
methods.
Occasion (within-cell replicate)
One of several ratings the same rater gives the
same subject. A design with occasions lets icc()
separate the subject-by-rater
interaction from pure error. glance() reports the
per-cell count in its n_o column. The printed report spells
that same count on its design line, as
N cells x N replicates. Where n_o is
NA because the cells hold unequal counts, that slot reads
NA too. Only a fit that splits replicates carries the slot
at all. A multilevel fit’s design line reports subjects and clusters
instead, and a fit that splits none reports observations or ratings, so
read n_o for those. See D-studies and within-cell
replicates.
occasions vs. n_o
Two near-identical names for different quantities. On an
icc() fit, tidy()’s occasions
column is the per-rater occasion divisor that row’s
coefficient applies to pure error. It reads 1 on every row that averages
no occasions, which includes every row whose error set carries no
pure-error term to average. It reads the fitted per-cell count where the
row does average, and NA on a fit that splits no
within-cell replicates. On a d_study() projection the same
column reports the count each row is projected at instead,
which ?d_study states. glance()’s
n_o counts the occasions observed per cell
in the design that was fitted, and is NA under the
condition ?icc states for it. So on a design with three
ratings per cell, a fit reporting both settings shows
occasions 1 and 3 down its rows, while n_o is
3 for the fit as a whole. Note that occasions is not a
count of the ratings a coefficient averages: an occasion-averaged
ICC(A,k) over four raters at occasions 3 is
the reliability of a mean of twelve ratings. A ragged replicate fit is
the case that separates the two columns most sharply: n_o
reads NA there while occasions still reads
1.
One-way vs. two-way
The model argument. A two-way design
has every subject rated by the same raters, so a rater main
effect can be estimated and either counted as error (agreement) or set
aside (consistency). A one-way design has each subject
rated by possibly different raters, so rater identity is not
modeled and only an agreement-style ICC(1) /
ICC(k) is defined.
Parametric bootstrap
An alternative confidence-interval method
(ci_method = "bootstrap"): simulate new response vectors
from the fitted model, refit each one, and take percentile quantiles of
the recomputed ICCs. It does not rely on the asymptotic-normal
approximation the Monte-Carlo method uses, at the cost of a full refit
per resample. See Confidence-interval
methods.
Posterior mode (MAP)
The point estimate the Bayesian brms engine reports: the peak (mode) of the posterior distribution of the ICC, the maximum a posteriori value. On a small, right-skewed posterior it can sit below the mixed-model REML estimate. Its interval is a credible interval. See Estimation engines.
Prior
In the Bayesian brms engine, the distribution placed
on each variance component before seeing the data. Here that is
a weakly-informative half-t(4, 0, 1) on every standard
deviation (ten Hove et al. 2020), the sourced prior every coverage
result depends on. Overriding it (prior =) voids those
guarantees, so icc() warns. See Estimation
engines.
REML
Restricted maximum likelihood: the standard method the mixed-model engines (glmmTMB, lme4) use to estimate variance components. It corrects the downward bias that ordinary maximum likelihood has when estimating variances, which matters for the small samples common in reliability studies.
Single-unit ICC: ICC(*,1)
The reliability of a single rater’s score. Its
counterpart, the average-unit ICC
ICC(*,k), is the reliability of the mean of k
raters and is always at least as high.
Subject level vs. cluster level
In a multilevel design (subjects nested in clusters,
such as pupils in classrooms), two reliabilities are defined. The
subject level asks how reliably raters distinguish
subjects within a cluster. The cluster level
asks how reliably they distinguish cluster means.
icc() reports both from one fit. See Multilevel
designs.
Transformed bootstrap-t
An opt-in confidence-interval method for the one-way design, balanced
or unbalanced (ci_method = "npbootstrap"; Ukoumunne et
al. 2003). It resamples whole subjects with replacement, stabilizes the
variance with a log-F transform, studentizes, and back-transforms the
endpoints, so it takes a seed and
boot_samples. Robust at the zero-variance boundary and to
non-normal subject effects. See Confidence-interval
methods.
Variance component
A share of the total variation in the scores traced to one source. It says how much comes from real differences between subjects, from some raters scoring higher than others, from the subject-by-rater interaction, and from residual error. Every ICC is a ratio built from these components: signal variance over signal-plus-error variance.
Zero-variance boundary
A variance component cannot be negative, so its estimate can land exactly at zero, the edge of the allowed range. Ordinary interval formulas misbehave there (they can run below zero or collapse). The Monte-Carlo default is boundary-aware: it works on a scale where zero is reachable without breaking, which is one reason glmmTMB is the recommended engine. See Confidence-interval methods.
References
Burch, B. D. (2011). Assessing the performance of normal-based and REML-based confidence intervals for the intraclass correlation coefficient. Computational Statistics and Data Analysis, 55, 1018–1028.
Jorgensen, T. D. (2021). How to estimate absolute-error components in structural equation models of generalizability theory. Psych, 3(2), 113–133.
McGraw, K. O., & Wong, S. P. (1996). Forming inferences about some intraclass correlation coefficients. Psychological Methods, 1(1), 30–46.
Searle, S. R. (1971). Linear Models. Wiley.
ten Hove, D., Jorgensen, T. D., & van der Ark, L. A. (2020). On the usefulness of interrater reliability coefficients. In M. Wiberg et al. (Eds.), Quantitative Psychology (pp. 67–75). Springer.
ten Hove, D., Jorgensen, T. D., & van der Ark, L. A. (2022). Interrater reliability for multilevel data: A generalizability theory approach. Psychological Methods, 27(4), 650–666.
Ukoumunne, O. C., Davison, A. C., Gulliford, M. C., & Chinn, S. (2003). Non-parametric bootstrap confidence intervals for the intraclass correlation coefficient. Statistics in Medicine, 22(24), 3805–3821.
Xiao, Y., & Liu, H. (2013). Modified profile likelihood approach for certain intraclass correlation coefficients. Computational Statistics, 28(5), 2241–2265.
