What an ICC tells you
Suppose several raters each score the same set of things: essays, patients, video clips. You want to know whether the scores can be trusted. If two raters watch the same clip, will they land on the same number? If you swapped in a different rater, would the result hold up? That is a question about interrater reliability. An intraclass correlation coefficient (ICC) is the number that answers it.
An ICC runs from 0 to 1. It reports the share of the variation in scores that reflects real differences between the things being rated, rather than disagreement or noise between raters. Near 1, almost all the spread in scores is genuine subject-to-subject difference. The raters barely disagree, and the ratings are highly reliable. Near 0, the raters are effectively adding noise. A score then tells you more about who happened to rate it than about the subject.
intraclass estimates that number by fitting a
mixed model, rather than from the classical ANOVA
mean-squares formulas older tools use. The model separates subject
variation from rater variation as variance components, each a
share of the total variation traced to one source. The two approaches
agree on clean, balanced data. The model-based one also handles missing
ratings, multiple designs, and honest confidence intervals. The Engines article covers the engine, the
software that does the fitting. This article walks the whole pipeline on
a small example: fit, estimate, interpret. New to any of the terms as
they come up? The Glossary defines
each one in a sentence.
The data
We use the classic Shrout & Fleiss (1979) example, shipped with
the package as ratings. It has 6 subjects each rated by the
same 4 raters, one rating per cell. The design is
two-way: each rater is tracked across the subjects they
score. intraclass wants long format: one
rating per row, with columns for the subject, the rater, and the
score.
head(ratings)
#> subject rater score
#> 1 1 1 9
#> 2 2 1 6
#> 3 3 1 8
#> 4 4 1 7
#> 5 5 1 10
#> 6 6 1 6Fit
Call icc() with the data and the three columns
(unquoted). With the defaults, a single two-way random fit reports every
defined formulation. It reports two kinds of coefficient.
Absolute agreement asks whether raters give the same
score. Consistency asks only that raters agree apart
from a constant offset per rater. Each kind comes in a single-rater
form, the reliability of one rater’s score, and an average-rater form,
the reliability of a mean of several raters. Those four are
ICC(A,1), ICC(A,k), ICC(C,1), and
ICC(C,k). The next sections unpack the labels. They are
grouped by error definition in the printout. We set a seed
so the confidence interval is reproducible.
fit <- icc(ratings, score, subject, rater, seed = 2024)
fit
#> ── Intraclass correlation: two-way random, absolute agreement & consistency ────
#> Subjects: 6 | Raters: 4 (random) | Observations: 24 of 24 cells (complete)
#> Engine: glmmTMB (REML) | CI: 95% montecarlo (10000 draws)
#>
#> index estimate 95% CI
#> Absolute agreement
#> ICC(A,1) 0.290 [0.050, 0.713]
#> ICC(A,k) 0.620 [0.173, 0.909]
#> Consistency
#> ICC(C,1) 0.715 [0.343, 0.924]
#> ICC(C,k) 0.909 [0.676, 0.980]
#>
#> Variance components: subject 2.556, rater 5.244, residual 1.019
#> Shrout & Fleiss equivalent: ICC(A,1) = ICC(2,1), ICC(A,k) = ICC(2,k)Read the result with the tidy verbs
tidy() returns one row per coefficient, with the
estimate and confidence interval. glance() returns a
one-row model summary, including the variance components.
tidy(fit)
#> # A tibble: 4 × 11
#> term occasions type level sf_index estimate std.error conf.low conf.high
#> <chr> <dbl> <chr> <chr> <chr> <dbl> <dbl> <dbl> <dbl>
#> 1 ICC(A,1) NA agree… NA ICC(2,1) 0.290 0.180 0.0498 0.713
#> 2 ICC(A,k) NA agree… NA ICC(2,k) 0.620 0.201 0.173 0.909
#> 3 ICC(C,1) NA consi… NA NA 0.715 0.155 0.343 0.924
#> 4 ICC(C,k) NA consi… NA NA 0.909 0.0810 0.676 0.980
#> # ℹ 2 more variables: conf.level <dbl>, method <chr>
glance(fit)
#> # A tibble: 1 × 24
#> n_subjects n_raters n_clusters n_obs n_cells balanced raters replicates
#> <int> <int> <int> <int> <int> <lgl> <chr> <lgl>
#> 1 6 4 NA 24 24 TRUE random FALSE
#> # ℹ 16 more variables: multilevel <lgl>, ml_design <chr>, k_eff <dbl>,
#> # k_c_eff <dbl>, var_cluster <dbl>, var_subject <dbl>, var_rater <dbl>,
#> # var_cluster_rater <dbl>, var_subject_rater <dbl>, var_residual <dbl>,
#> # n_o <int>, engine <chr>, ci_method <chr>, conf.level <dbl>, rhat <dbl>,
#> # ess_bulk <dbl>summary() is the other reading of the same fit. It
reprints the report above, then appends a short interpretive
note for each error definition present. It adds one more note,
about what a single rating per cell cannot separate. Nothing is
recomputed: the notes are read off the design, so they track whatever
you fitted:
summary(fit)
#> ── Intraclass correlation: two-way random, absolute agreement & consistency ────
#> Subjects: 6 | Raters: 4 (random) | Observations: 24 of 24 cells (complete)
#> Engine: glmmTMB (REML) | CI: 95% montecarlo (10000 draws)
#>
#> index estimate 95% CI
#> Absolute agreement
#> ICC(A,1) 0.290 [0.050, 0.713]
#> ICC(A,k) 0.620 [0.173, 0.909]
#> Consistency
#> ICC(C,1) 0.715 [0.343, 0.924]
#> ICC(C,k) 0.909 [0.676, 0.980]
#>
#> Variance components: subject 2.556, rater 5.244, residual 1.019
#> Shrout & Fleiss equivalent: ICC(A,1) = ICC(2,1), ICC(A,k) = ICC(2,k)
#>
#> Absolute agreement counts the rater main effect (systematic differences in rater level) as error.
#> Consistency ignores the rater main effect (systematic differences in rater level). Only relative standing counts.
#> A single rating per cell confounds the subject-by-rater interaction with
#> residual error.Interpret
Single vs. the average of several raters.
ICC(A,1) = 0.29 is the reliability of a single
rater: how much you can trust one person’s score. ICC(A,k)
= 0.62 is the reliability of the mean of all 4 raters.
Averaging cancels out independent rater noise, so the mean is always
more reliable than one rater alone. This is the
Spearman–Brown relationship: more raters, higher
reliability, with diminishing returns. Report ICC(A,k) when
the averaged score is what you will actually use. Report
ICC(A,1) when downstream users see one rater’s
judgment.
Absolute agreement vs. rank order. These are absolute-agreement coefficients. If one rater scores consistently higher than another, that systematic gap counts as error. Here the raters differ sharply in average level, so the agreement ICC is low. That is a signal that the rating procedure has a level problem worth fixing. If you only care that raters rank subjects the same way, ask for consistency instead, described below.
Is this a good ICC?
There is no universal cutoff. Two widely cited rules of thumb give a rough vocabulary. Koo & Li (2016) propose:
| ICC | Label |
|---|---|
| < 0.50 | poor |
| 0.50–0.75 | moderate |
| 0.75–0.90 | good |
| > 0.90 | excellent |
An older scheme, Cicchetti (1994), draws its lines a little differently: < 0.40 poor, 0.40–0.59 fair, 0.60–0.74 good, 0.75–1.00 excellent. Treat these as conventions, not laws. The bar that matters depends on the stakes of your decision, and on which ICC you are reading. An agreement coefficient and a consistency coefficient on the same data are not comparable to the same cutoff.
Most important, and this is Koo & Li’s own recommendation:
judge the confidence interval, not just the point
estimate. Here the four-rater mean ICC(A,k) = 0.62
reads as “moderate”. Its 95% interval runs from 0.17 to 0.91, from
“poor” all the way to “excellent.” With only six subjects, the data
simply cannot pin the reliability down to one band. So
icc() never returns a point estimate without an interval,
and never labels a result for you: the honest summary is the whole
interval.
About the confidence interval
The interval icc() reports is a Monte-Carlo
interval, built by drawing parameter values from the fitted model’s
uncertainty, and it is boundary-aware.
In plain terms: the default interval does not rely on a textbook formula
that misbehaves when a variance is near zero. That is a common
situation: raters who barely differ push the rater variance to its zero
boundary. Instead icc() simulates many plausible parameter
sets from the fitted model, and reads the interval off the resulting
spread of ICCs. This keeps the interval well-behaved right at that
boundary. The Interval
methods article covers how it works and the alternatives. One
alternative is a parametric bootstrap, which refits the model on
simulated data many times. Another is a Bayesian credible interval,
which holds the share of the posterior probability that the confidence
level sets.
Consistency instead of agreement
Suppose a constant per-rater offset is acceptable: you care only that
raters rank subjects the same way, not that they land on the
same number. Then ask for consistency with
type = "consistency". It drops the rater main effect from
the error, so it is never smaller than the agreement coefficient:
icc(ratings, score, subject, rater, type = "consistency", seed = 2024)
#> ── Intraclass correlation: two-way random, consistency ─────────────────────────
#> Subjects: 6 | Raters: 4 (random) | Observations: 24 of 24 cells (complete)
#> Engine: glmmTMB (REML) | CI: 95% montecarlo (10000 draws)
#>
#> index estimate 95% CI
#> ICC(C,1) 0.715 [0.343, 0.924]
#> ICC(C,k) 0.909 [0.676, 0.980]
#>
#> Variance components: subject 2.556, rater 5.244, residual 1.019The gap between the two (here ICC(C,1) = 0.71
vs. ICC(A,1) = 0.29) is a direct read-out of how much
systematic rater-level difference is present.
By default raters are treated as random: a sample
you want to generalize beyond, so that your reliability claim covers
raters you did not use. Raters are fixed when the
observed raters are the whole population of interest. If yours are, you
can pass raters = "fixed", the classic
ICC(3,1). But icc() will warn: random is the
recommended default for interrater reliability, and on balanced data the
number is identical anyway.
Which ICC do I want?
Several choices define which ICC is correct for your study.
Are the raters crossed, so that the same set judges everyone, or
interchangeable (model = "oneway")? Absolute agreement or
consistency? A single rater or the average? Fixed or random raters? A
complete or an incomplete design? The Choosing an ICC article
(vignette("choosing-an-icc")) walks that decision step by
step.
References
Cicchetti, D. V. (1994). Guidelines, criteria, and rules of thumb for evaluating normed and standardized assessment instruments in psychology. Psychological Assessment, 6(4), 284–290.
Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155–163.
Shrout, P. E., & Fleiss, J. L. (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428.
