Skip to contents

intraclass (development version)

  • The package now ships a citation entry (citation("intraclass")), and the source repository gains a CONTRIBUTING.md that says how to report a problem, ask for help, and propose a change. The README carries a Citation section and links to the contributing guide and to the companion software paper’s repository, https://github.com/jmgirard/intraclass-paper.

  • Four simulated teaching datasets now ship with the package: school (pupils nested in classrooms), school_incomplete (the same design with a fifth of the ratings dropped), ratings_replicates (three ratings per subject-by-rater cell) and ratings_twoway (20 subjects by 4 raters). The articles and the README load them instead of simulating data inline, so the examples can be run from a fresh session with no setup. Each has a reference page that names the generating script and seed. The values the articles report do not change. The README’s multilevel example, which used its own smaller simulation before, now shows the school results.

Documentation

  • The package’s own errors, warnings, notes and install prompts now use plainer English. None uses a dash as punctuation, and no sentence runs past 25 words. Long sentences are split, and semicolons become full stops except between references. Which condition fires, and when, is unchanged. Nothing computed changes.

  • The questions and advice that choose_icc() prints, and the notes that print() and summary() add to an icc() fit, now use plainer English. None uses a dash as punctuation, and no sentence runs past 25 words. Which note prints, and when, is unchanged. Five definitions in the Glossary article were imprecise or said less than the package does, and they are corrected. Every page that quoted them now carries the new text. Four error messages from icc() and d_study() no longer cite internal record numbers. Nothing computed changes.

  • Four articles now show their comparison tables as gt tables with a title, plain column headings and rounded numbers. The four are Comparison with other packages, Estimation engines (an engine is the software that does the fitting), Confidence-interval methods and Choosing an ICC. The code that assembles each table is hidden. The calls that fit the models stay visible. The one exception is the psych and irr comparison: there, one hidden chunk fits the models and builds the table. gt joins Suggests:. Without it, those tables are skipped. The sentences that disagreed with their tables are corrected. The packages agree to within 0.00001, not exactly. In the Choosing an ICC incomplete-data example, the random-rater interval is wider only for the agreement coefficients, which ask whether raters give the same score.

  • The reference manual pages for icc(), d_study(), choose_icc(), ratings and ratings_incomplete are rewritten in plainer English. No sentence runs past 25 words, except seven on the icc() and d_study() pages that carry clauses the test suite pins verbatim. Glossary terms are defined in plain words where the pages first use them. The icc() page is now titled “Intraclass correlation coefficients for interrater reliability” and the d_study() page “Project reliability to other numbers of raters”. Nothing computed changes.

  • The Getting started, Choosing an ICC and Glossary articles and the README are rewritten in plainer English. No sentence runs past 25 words, except two glossary sentences that carry clauses the test suite pins verbatim. Each glossary term is defined in plain words at its first use in Getting started, Choosing an ICC and the README. The glossary’s entries each open with a one-sentence definition. Nothing computed changes.

  • Five more articles are rewritten in plainer English. A D-study projects the fitted variance components to other rater or occasion counts, each a share of the total variation traced to one source. One of the five covers D-studies and the within-cell replicate, one of several ratings by the same rater of the same subject. The five are Estimation engines, Comparison with other packages, D-studies and within-cell replicates, Multilevel designs and Confidence-interval methods. No sentence runs past 25 words, except two in the interval-methods article that carry clauses the test suite pins verbatim. Each glossary term is defined in plain words or linked to the glossary at its first use in each article. In the interval-methods article every ci_method section now opens with a paragraph on what the method is for and when to choose it. Nothing computed changes.

intraclass 0.1.0

CRAN release: 2026-09-10

First public release.

intraclass estimates interrater-reliability intraclass correlation coefficients (ICCs) within the generalizability-theory framework. Variance components come from a linear mixed model rather than from classical ANOVA mean squares. A point estimate is never reported without an interval, and ci_method selects which interval that is. The package requires R 4.5.0 or newer.

What ships

  • icc() fits the model and reports the ICC family the design defines. The type argument picks absolute agreement or consistency, where consistency means raters agree apart from a constant offset per rater. The unit argument picks single or average. The raters argument picks random or fixed raters, where fixed means the observed raters are the whole population of interest. The model argument picks one-way or two-way, where two-way means each rater is tracked across the subjects they score. See Getting started and ?icc.
  • Data need not form a complete, balanced grid. There is support for imbalanced, incomplete, and multilevel (nested) designs. That covers unequal ratings per subject, missing ratings, and subjects nested in a higher-level unit such as a classroom or clinic. A cluster column on a two-way design switches on the multilevel ICC. In a multilevel design the subject level is reliability within a cluster, and the cluster level is reliability of cluster means. When the same raters span every cluster, the column also adds the cluster level. See Multilevel designs: subject and cluster level and ?icc for the layouts and where it refuses.
  • Rating a subject-by-rater cell more than once gives a within-cell replicate design. On a two-way random design icc() splits the single-rating residual into a subject-by-rater interaction and pure error, and occasions reports the reliability of one rating or, on balanced replicates, of the mean. See D-studies and within-cell replicates and ?icc for the designs replicates support and where they refuse.
  • d_study() projects a fitted reliability to other numbers of raters (m) or occasions (n_o), with a plot() and ggplot2::autoplot() curve. D-studies and within-cell replicates gives the designs each projection supports and where it refuses.
  • choose_icc() recommends which coefficient or coefficients to report, explains the reasoning, and prints the icc() call to run. Its type, unit and level questions each take "both", which asks for the pair rather than making you choose one. It gives advice only. See Choosing an ICC.
  • tidy() and glance() return tidy summaries of a fit or a projection. print(), format(), plot() and ggplot2::autoplot() methods are provided for both classes, and summary() for an icc() fit. ggplot2 is a Suggests dependency.
  • Datasets ratings and ratings_incomplete, used throughout the documentation.
  • Eight articles: Getting started, Choosing an ICC, Multilevel designs: subject and cluster level, Estimation engines, Confidence-interval methods, D-studies and within-cell replicates, Glossary, and Comparison with other packages.

Engines

  • The default engine is glmmTMB, the package’s only hard engine dependency. engine = "lme4", engine = "lavaan" and engine = "brms" are selectable. Which designs each engine covers, and where it refuses, is documented in Estimation engines and ?icc.
  • The brms engine uses a prior, the distribution placed on a parameter before seeing the data. It fits both random- and fixed-rater models under a sourced half-t(4, 0, 1) prior on every random-effect standard deviation. Its point estimate is the posterior mode, the peak of the posterior distribution. Its interval is a percentile credible interval, which holds the share of the posterior probability that the confidence level sets. Because the interval comes from the posterior draws, ci_method = "posterior" is forced. Supplying a custom prior is a deliberate deviation: icc() warns, and the coverage results this package reports no longer apply.
  • The lme4 package itself is already on your library path after a plain install, because glmmTMB lists it in its own Imports, but the lme4 engine also needs merDeriv, which does not arrive. merDeriv, lavaan and brms sit in this package’s Suggests, and a plain install fetches none of the three.

Confidence intervals

  • ci_method selects the interval: "montecarlo" (the default), "bootstrap", "npbootstrap", "searle", "burch", "mpl", and "posterior" under brms. The default is the Monte-Carlo interval, built by drawing parameter values from the fitted model’s uncertainty. Where a method does not apply to the design, the call aborts with a classed error. Confidence-interval methods compares them; ?icc gives the per-method conditions.
  • Interval coverage was studied by simulation before release, and the limitations those studies found are documented rather than smoothed over. The Monte-Carlo default under-covers when the subject effects are strongly skewed or heavy-tailed, worst 0.6725 at chi-square(1) subject effects with a true ICC of 0.6, 50 subjects and 5 raters. "burch" is no remedy there (worst 0.6655).
  • "searle" and "burch" are the two classical closed forms, and the three grids below measure their widths: the smaller grid’s 16 cells and the larger grid’s 64 cells draw only the subject effects from the non-normal family. The third draws the residual from the same family as the subject effect.
  • Which of the two closed forms gives the narrower interval is conditional. On both grids that vary only the subject effect, "burch" is the narrower of the two in 16 of 16 cells of the smaller grid and 59 of 64 cells of the larger grid.
  • That width margin holds much the same up to a true ICC of 0.3 rather than shrinking as the true ICC rises (on the larger grid; the smaller grid’s margin does shrink across its levels). The "burch" width advantage then collapses to near parity at a true ICC of 0.6, on the one grid reaching that value. It also shrinks steadily as the subject count grows, measured at 5 raters.
  • What "burch" does against "searle" also depends on the residual, and the three grids measure that: the two grids that vary only the subject effect put it narrower nearly everywhere, while the third, which draws the residual from the same family as the subject effect, puts it wider at every symmetric heavy-tailed family measured (a median width ratio of 1.2963 at t(5) with 100 subjects) and narrower at every lighter-tailed one, the normal included.
  • The ci_method = "mpl" documentation states the interpolation evidence behind off-node subject counts. The correction constant is calibrated at subject-count nodes and linearly interpolated between them. The interpolated path is coverage-validated at each supported confidence level, at the default 0.95 by three off-node cells, each clearing its pre-registered coverage floor. The documentation also says what that validation does not establish (interpolated values are validated at a handful of geometries, not calibrated, and the interval’s asymmetry direction is not uniform across rater counts). Every universal or negative claim this documentation makes about the validated cells is settled mechanically, in CI, against the committed coverage fixture (data-raw/check-mpl-doc-claims.py).

When a call fails

  • Errors are classed and name the quantities that failed the check, rather than asserting a cause. Where an interval method aborts on degenerate data, the error names another ci_method only after running it on your data and confirming it returns a usable interval.