A plain-language reference for the terms that recur across these articles. Each entry defines the idea once, and the other articles link here rather than re-explaining it. Terms are listed alphabetically. Nothing here is new: it is the vocabulary the Getting started and Choosing an ICC guides use, gathered in one place.
Absolute agreement
One of the two types of ICC: raters give the same score.
Two raters who rank subjects identically but sit a full point apart do
not agree in this sense. It counts systematic rater
differences (the rater variance) as error. Contrast
consistency. In generalizability theory the
absolute-agreement ICC is the dependability
coefficient.
Average-unit ICC: ICC(*,k)
The reliability of a mean of several raters, k of them,
rather than of one rater. Averaging cancels part of the error, so
ICC(*,k) is always at least as high as the single-rater
ICC(*,1). The k is the number of raters whose
average you will actually use. On incomplete data it becomes the
effective number of ratings. See single-unit
ICC for its counterpart.
Burch interval
An opt-in confidence interval for the balanced one-way design
(ci_method = "burch"). It is closed-form, meaning computed
from a formula with no resampling. It gives REML-based limits with a
kurtosis adjustment (Burch 2011). Kurtosis is tail weight: how often a
distribution produces extreme values. So the Burch interval’s width
tracks the tail weight of the data, where the exact-F
interval leans on normality instead. Tracking tail weight is
not the same as widening. On the two grids this package measured that
vary only the subject effect, it was the narrower in nearly every cell.
How much narrower is conditional, and not in the direction one might
guess. "burch"’s width margin holds much the same up to a
true ICC of 0.3 rather than shrinking as the true ICC rises (on the
larger grid; the smaller grid’s margin does shrink across its levels).
The "burch" width margin then collapses to near parity at a
true ICC of 0.6, on the one grid reaching that value. It shrinks
steadily as the subject count grows, measured at 5 raters. Every cell
favouring the exact-F interval sits on the grid where the margin reaches
parity. Burch reports the reverse for symmetric heavy-tailed data with
non-normal errors, and a third grid now measures that case. What
"burch" does against "searle" depends on what
the residual is drawn from. The three grids now measure that: the two
grids that vary only the subject effect put it narrower nearly
everywhere, while the third, which draws the residual from the same
family as the subject effect, puts it wider at every symmetric
heavy-tailed family measured (a median width ratio of 1.2963 at t(5)
with 100 subjects) and narrower at every lighter-tailed one, the normal
included. So neither interval is reliably the tighter one. That
adjustment is not a remedy for heavy tails: on strongly skewed subject
effects "burch" under-covers about as badly as the default.
See When the
default under-covers. Deterministic: no resamples, no seed. See
Confidence-interval
methods.
Confidence interval vs. credible interval
A confidence interval is a range built with many
hypothetical repetitions of the study in mind. Across them, the chosen
share (95% by default) of such ranges would contain the true ICC. The
frequentist engines output one. A credible interval is
a range that holds the share of the posterior probability that the
confidence level sets. That level is the conf_level
argument, 0.95 by default. The Bayesian brms engine outputs
one. At the default level you can say directly “there is a 95% chance
the ICC lies in here, given the data and the prior.” They answer subtly
different questions. See Confidence-interval
methods.
Conflated ICC
The single-level ICC that ignores clustering, such as pupils in
classrooms or patients in clinics. This quantity is ten Hove et al.’s
(2022) Equation 14. It folds the between-cluster and within-cluster
variation into one “true score”. It is biased for both the subject-level
and cluster-level questions. icc() can report it
(level = "conflated") purely as a diagnostic
contrast, to show the cost of ignoring the structure. It is
never a number to report. See Multilevel
designs.
Connectedness (identification)
A design is connected when the raters and subjects
form one linked web. The web must be tight enough that the model can
separate a subject effect from a rater effect. With enough missing cells
a design can split into disconnected islands. Then the variance
components are not identified, meaning no method can
estimate them. icc() checks this and aborts loudly rather
than return a number that is not estimable.
Consistency
The other type of ICC: raters agree apart from a
constant offset per rater. Two raters who agree on the ordering but
differ by a fixed point still count as perfectly consistent. It leaves
the rater main effect out of the error term. Contrast absolute
agreement.
Dependability coefficient
The generalizability-theory name for the absolute-agreement ICC, written . Projecting it to a different number of raters (a D-study) is a change of the averaging divisor. See D-studies and within-cell replicates.
D-study (decision study)
A forward-looking projection of the fitted variance components to other rater or occasion counts. It asks how reliable the mean of some other number of raters (or occasions) would be. It reuses the variance components you already estimated, with no refitting. It answers “how many raters do I need?”. See D-studies and within-cell replicates.
Effective number of ratings: k_eff
The number of ratings the average is really over, on
incomplete data. There, subjects are rated different
numbers of times, so there is no single k to average over.
k_eff is the harmonic mean of the
per-subject rating counts. It is the divisor icc() uses for
ICC(*,k) on ragged data. It is always at or below the full
panel size. The report names it, so the divisor is never a black
box.
Engine
The backend that fits the model: the software icc() uses
to estimate the variance components, chosen with the engine
argument. The choices are glmmTMB (the default mixed
model), lme4 (an alternate mixed-model solver),
lavaan (a structural-equation formulation), and
brms (a Bayesian fit). Some engines are just a
different solver for the same estimator. Others compute a genuinely
different estimator, though one that agrees in large samples. See Estimation engines.
Estimand
The true quantity you are trying to estimate: the target the ICC is aiming at. The word matters because “the ICC” is not one number. Agreement and consistency, single and average, subject level and cluster level are different estimands. Picking the coefficient is really picking which one answers your question. See Choosing an ICC.
Exact-F interval
An opt-in confidence interval for the balanced one-way design
(ci_method = "searle") that assumes normal data. It is
closed-form: it inverts the exact-F pivot of the one-way ANOVA (Searle
1971, Ch. 9 Table 9.14; the McGraw & Wong 1996 Table 7 limits).
Exact under normality. It is best-calibrated when the data are
approximately normal, and it is deterministic. Exact is not the same as
shortest. See the Burch interval above for what this
package measured about their relative widths. See Confidence-interval
methods.
FIML
Full-information maximum likelihood: a fitting technique that uses every observed value rather than dropping incomplete cases. The lavaan (SEM) engine uses it to fit incomplete data. See Estimation engines.
Finite-population rater variance: θ²_r
The spread of just the observed raters, used when raters are treated as fixed. Fixed means the observed raters are the whole population of interest. Then the “rater variance” is computed as a bias-corrected finite-population quantity (McGraw & Wong’s Case 3A). It is not an estimate of a wider rater universe. On balanced data it equals the random-rater variance. Under imbalance it differs. See fixed vs. random raters.
Fixed vs. random raters
The raters argument. Random raters (the
recommended default) treat the raters you used as a sample from a larger
pool. So the reliability generalizes to new raters drawn from
that pool. Fixed raters treat the observed raters as
the entire population of interest. So the reliability speaks only to
these raters. The choice changes the rater term from a
random-sample variance to the finite-population rater
variance.
Harmonic mean
An average that leans toward the smaller values: the reciprocal of the mean of the reciprocals. It is the right average for the effective number of ratings. Reliability depends on the rate of information per subject, which the harmonic mean captures.
Indicator-mean estimator
How the lavaan (SEM) engine recovers the rater variance for absolute agreement. In lavaan’s formulation a rater is a single column with no random effect. So its variance is read from the spread of the estimated column (indicator) means (Jorgensen 2021). It is a genuinely different estimator than the mixed model’s random effect. The two agree in large samples but can differ modestly on small designs. See Estimation engines.
Modified profile likelihood
An opt-in deterministic confidence interval for the balanced,
complete two-way random absolute-agreement design
(ci_method = "mpl"; Xiao & Liu 2013). It profiles the
likelihood in the ICC, with a calibrated small-sample correction.
Profiling means treating the ICC as the one unknown and maximizing over
the rest. It returns an interval at the near-zero boundary where the
Monte-Carlo default aborts. Available at conf_level 0.90,
0.95, and 0.99, and deliberately conservative. See Confidence-interval
methods.
Monte-Carlo interval
The default confidence-interval method
(ci_method = "montecarlo"), built by drawing parameter
values from the fitted model’s uncertainty. It draws many parameter
vectors from the fitted model’s estimated covariance, on a scale that
respects the zero-variance boundary. It simulates no new data. It
recomputes the ICC for each draw, and takes the 2.5% and 97.5%
quantiles. Fast and boundary-aware. It does assume the fitted parameters
are approximately normally distributed around the truth. It under-covers
when the subject effects are strongly skewed or heavy-tailed. See When the
default under-covers. See also Confidence-interval
methods.
Occasion (within-cell replicate)
One of several ratings by the same rater of the same subject. A
design with occasions lets icc() separate the subject-by-rater interaction from pure
error. glance() reports the per-cell count in its
n_o column. The printed report spells that same count on
its design line, as N cells x N replicates. Where
n_o is NA because the cells hold unequal
counts, that slot reads NA too. Only a fit that splits
replicates carries the slot at all. A multilevel fit’s design line
reports subjects and clusters instead. A fit that splits none reports
observations or ratings. So read n_o for those. See D-studies and within-cell
replicates.
occasions vs. n_o
Two near-identical names for different quantities. On an
icc() fit, tidy()’s occasions
column is the per-rater occasion divisor that row’s
coefficient applies to pure error. It reads 1 on every row that averages
no occasions. That includes every row whose error set carries no
pure-error term to average. It reads the fitted per-cell count where the
row does average. It reads NA on a fit that splits no
within-cell replicates. On a d_study() projection the same
column reports the count each row is projected at instead,
which ?d_study states. glance()’s
n_o counts the occasions observed per cell
in the design that was fitted. It is NA under the condition
?icc states for it. So take a design with three ratings per
cell. A fit reporting both settings shows occasions 1 and 3
down its rows, while n_o is 3 for the fit as a whole. Note
that occasions is not a count of the ratings a coefficient
averages. An occasion-averaged ICC(A,k) over four raters at
occasions 3 is the reliability of a mean of twelve ratings.
A ragged replicate fit separates the two columns most sharply:
n_o reads NA there while
occasions still reads 1.
One-way vs. two-way
The model argument. In a two-way design
each rater is tracked across the subjects they score, whether every
rater scores every subject or not. So a rater main effect can be
estimated, and either counted as error (agreement) or set aside
(consistency). A one-way design has each subject rated
by possibly different raters. So rater identity is not modeled,
and only an agreement-style ICC(1) / ICC(k) is
defined.
Parametric bootstrap
A confidence-interval method (ci_method = "bootstrap")
that refits the model on simulated data many times. It simulates new
response vectors from the fitted model, refits each one, and takes
percentile quantiles of the recomputed ICCs. It does not rely on the
large-sample normal approximation the Monte-Carlo method uses. The cost
is a full refit per resample. See Confidence-interval
methods.
Posterior mode (MAP)
The point estimate the Bayesian brms engine reports: the peak of the posterior distribution of the ICC, the maximum a posteriori value. The posterior is the distribution of the ICC after seeing the data. On a small, right-skewed posterior it can sit below the mixed-model REML estimate. Its interval is a credible interval. See Estimation engines.
Prior
In the Bayesian brms engine, the distribution placed
on a parameter before seeing the data, here on each variance component.
Here that is a weakly-informative half-t(4, 0, 1) on every
standard deviation (ten Hove et al. 2020). It is the sourced prior every
coverage result depends on. Overriding it (prior =) voids
those guarantees, so icc() warns. See Estimation
engines.
REML
Restricted maximum likelihood: a way to estimate variances that corrects maximum likelihood’s downward bias. It is the standard method the mixed-model engines (glmmTMB, lme4) use to estimate variance components. Ordinary maximum likelihood underestimates variances, which matters for the small samples common in reliability studies. REML corrects that.
Single-unit ICC: ICC(*,1)
The reliability of one rater’s score. Its counterpart, the
average-unit ICC ICC(*,k), is the
reliability of the mean of k raters and is always at least
as high.
Subject level vs. cluster level
Two reliabilities defined in a multilevel design,
where subjects are nested in clusters such as pupils in classrooms. Here
the subject level is reliability within a cluster, and the cluster level
is reliability of cluster means. The subject level asks
how reliably raters distinguish subjects within a cluster. The
cluster level asks how reliably raters distinguish
cluster means. icc() reports both from one fit.
See Multilevel
designs.
Transformed bootstrap-t
An opt-in confidence-interval method for the one-way design, balanced
or unbalanced (ci_method = "npbootstrap"; Ukoumunne et
al. 2003). It resamples whole subjects with replacement. It stabilizes
the variance with a log-F transform. It then studentizes, meaning it
scales each resampled estimate, on the transformed scale, by its own
standard error, and back-transforms the endpoints. So it takes a
seed and boot_samples. Robust at the
zero-variance boundary and to non-normal subject effects. See Confidence-interval
methods.
Variance component
A share of the total variation in the scores traced to one source. It says how much comes from real differences between subjects, from some raters scoring higher than others, from the subject-by-rater interaction, and from residual error. Every ICC is a ratio built from these components: signal variance over signal-plus-error variance.
Zero-variance boundary
The edge of the allowed range, where a variance estimate lands exactly at zero. A variance component cannot be negative, so its estimate can land there. Ordinary interval formulas misbehave at that edge: they can run below zero or collapse. The Monte-Carlo default is boundary-aware. It works on a scale where zero is reachable without breaking, which is one reason glmmTMB is the recommended engine. See Confidence-interval methods.
References
Burch, B. D. (2011). Assessing the performance of normal-based and REML-based confidence intervals for the intraclass correlation coefficient. Computational Statistics and Data Analysis, 55, 1018–1028.
Jorgensen, T. D. (2021). How to estimate absolute-error components in structural equation models of generalizability theory. Psych, 3(2), 113–133.
McGraw, K. O., & Wong, S. P. (1996). Forming inferences about some intraclass correlation coefficients. Psychological Methods, 1(1), 30–46.
Searle, S. R. (1971). Linear Models. Wiley.
ten Hove, D., Jorgensen, T. D., & van der Ark, L. A. (2020). On the usefulness of interrater reliability coefficients. In M. Wiberg et al. (Eds.), Quantitative Psychology (pp. 67–75). Springer.
ten Hove, D., Jorgensen, T. D., & van der Ark, L. A. (2022). Interrater reliability for multilevel data: A generalizability theory approach. Psychological Methods, 27(4), 650–666.
Ukoumunne, O. C., Davison, A. C., Gulliford, M. C., & Chinn, S. (2003). Non-parametric bootstrap confidence intervals for the intraclass correlation coefficient. Statistics in Medicine, 22(24), 3805–3821.
Xiao, Y., & Liu, H. (2013). Modified profile likelihood approach for certain intraclass correlation coefficients. Computational Statistics, 28(5), 2241–2265.
