intraclass (development version)
The package now ships a citation entry (
citation("intraclass")), and the source repository gains aCONTRIBUTING.mdthat says how to report a problem, ask for help, and propose a change. The README carries a Citation section and links to the contributing guide and to the companion software paper’s repository, https://github.com/jmgirard/intraclass-paper.Four simulated teaching datasets now ship with the package:
school(pupils nested in classrooms),school_incomplete(the same design with a fifth of the ratings dropped),ratings_replicates(three ratings per subject-by-rater cell) andratings_twoway(20 subjects by 4 raters). The articles and the README load them instead of simulating data inline, so the examples can be run from a fresh session with no setup. Each has a reference page that names the generating script and seed. The values the articles report do not change. The README’s multilevel example, which used its own smaller simulation before, now shows theschoolresults.
Documentation
The package’s own errors, warnings, notes and install prompts now use plainer English. None uses a dash as punctuation, and no sentence runs past 25 words. Long sentences are split, and semicolons become full stops except between references. Which condition fires, and when, is unchanged. Nothing computed changes.
The questions and advice that
choose_icc()prints, and the notes thatprint()andsummary()add to anicc()fit, now use plainer English. None uses a dash as punctuation, and no sentence runs past 25 words. Which note prints, and when, is unchanged. Five definitions in the Glossary article were imprecise or said less than the package does, and they are corrected. Every page that quoted them now carries the new text. Four error messages fromicc()andd_study()no longer cite internal record numbers. Nothing computed changes.Four articles now show their comparison tables as
gttables with a title, plain column headings and rounded numbers. The four are Comparison with other packages, Estimation engines (an engine is the software that does the fitting), Confidence-interval methods and Choosing an ICC. The code that assembles each table is hidden. The calls that fit the models stay visible. The one exception is thepsychandirrcomparison: there, one hidden chunk fits the models and builds the table.gtjoinsSuggests:. Without it, those tables are skipped. The sentences that disagreed with their tables are corrected. The packages agree to within 0.00001, not exactly. In the Choosing an ICC incomplete-data example, the random-rater interval is wider only for the agreement coefficients, which ask whether raters give the same score.The reference manual pages for
icc(),d_study(),choose_icc(),ratingsandratings_incompleteare rewritten in plainer English. No sentence runs past 25 words, except seven on theicc()andd_study()pages that carry clauses the test suite pins verbatim. Glossary terms are defined in plain words where the pages first use them. Theicc()page is now titled “Intraclass correlation coefficients for interrater reliability” and thed_study()page “Project reliability to other numbers of raters”. Nothing computed changes.The Getting started, Choosing an ICC and Glossary articles and the README are rewritten in plainer English. No sentence runs past 25 words, except two glossary sentences that carry clauses the test suite pins verbatim. Each glossary term is defined in plain words at its first use in Getting started, Choosing an ICC and the README. The glossary’s entries each open with a one-sentence definition. Nothing computed changes.
Five more articles are rewritten in plainer English. A D-study projects the fitted variance components to other rater or occasion counts, each a share of the total variation traced to one source. One of the five covers D-studies and the within-cell replicate, one of several ratings by the same rater of the same subject. The five are Estimation engines, Comparison with other packages, D-studies and within-cell replicates, Multilevel designs and Confidence-interval methods. No sentence runs past 25 words, except two in the interval-methods article that carry clauses the test suite pins verbatim. Each glossary term is defined in plain words or linked to the glossary at its first use in each article. In the interval-methods article every
ci_methodsection now opens with a paragraph on what the method is for and when to choose it. Nothing computed changes.
intraclass 0.1.0
CRAN release: 2026-09-10
First public release.
intraclass estimates interrater-reliability intraclass correlation coefficients (ICCs) within the generalizability-theory framework. Variance components come from a linear mixed model rather than from classical ANOVA mean squares. A point estimate is never reported without an interval, and ci_method selects which interval that is. The package requires R 4.5.0 or newer.
What ships
-
icc()fits the model and reports the ICC family the design defines. Thetypeargument picks absolute agreement or consistency, where consistency means raters agree apart from a constant offset per rater. Theunitargument picks single or average. Theratersargument picks random or fixed raters, where fixed means the observed raters are the whole population of interest. Themodelargument picks one-way or two-way, where two-way means each rater is tracked across the subjects they score. See Getting started and?icc. - Data need not form a complete, balanced grid. There is support for imbalanced, incomplete, and multilevel (nested) designs. That covers unequal ratings per subject, missing ratings, and subjects nested in a higher-level unit such as a classroom or clinic. A
clustercolumn on a two-way design switches on the multilevel ICC. In a multilevel design the subject level is reliability within a cluster, and the cluster level is reliability of cluster means. When the same raters span every cluster, the column also adds the cluster level. See Multilevel designs: subject and cluster level and?iccfor the layouts and where it refuses. - Rating a subject-by-rater cell more than once gives a within-cell replicate design. On a two-way random design
icc()splits the single-rating residual into a subject-by-rater interaction and pure error, andoccasionsreports the reliability of one rating or, on balanced replicates, of the mean. See D-studies and within-cell replicates and?iccfor the designs replicates support and where they refuse. -
d_study()projects a fitted reliability to other numbers of raters (m) or occasions (n_o), with aplot()andggplot2::autoplot()curve. D-studies and within-cell replicates gives the designs each projection supports and where it refuses. -
choose_icc()recommends which coefficient or coefficients to report, explains the reasoning, and prints theicc()call to run. Itstype,unitandlevelquestions each take"both", which asks for the pair rather than making you choose one. It gives advice only. See Choosing an ICC. -
tidy()andglance()return tidy summaries of a fit or a projection.print(),format(),plot()andggplot2::autoplot()methods are provided for both classes, andsummary()for anicc()fit.ggplot2is aSuggestsdependency. - Datasets
ratingsandratings_incomplete, used throughout the documentation. - Eight articles: Getting started, Choosing an ICC, Multilevel designs: subject and cluster level, Estimation engines, Confidence-interval methods, D-studies and within-cell replicates, Glossary, and Comparison with other packages.
Engines
- The default engine is glmmTMB, the package’s only hard engine dependency.
engine = "lme4",engine = "lavaan"andengine = "brms"are selectable. Which designs each engine covers, and where it refuses, is documented in Estimation engines and?icc. - The brms engine uses a prior, the distribution placed on a parameter before seeing the data. It fits both random- and fixed-rater models under a sourced half-t(4, 0, 1) prior on every random-effect standard deviation. Its point estimate is the posterior mode, the peak of the posterior distribution. Its interval is a percentile credible interval, which holds the share of the posterior probability that the confidence level sets. Because the interval comes from the posterior draws,
ci_method = "posterior"is forced. Supplying a customprioris a deliberate deviation:icc()warns, and the coverage results this package reports no longer apply. - The
lme4package itself is already on your library path after a plain install, becauseglmmTMBlists it in its ownImports, but the lme4 engine also needs merDeriv, which does not arrive.merDeriv,lavaanandbrmssit in this package’sSuggests, and a plain install fetches none of the three.
Confidence intervals
-
ci_methodselects the interval:"montecarlo"(the default),"bootstrap","npbootstrap","searle","burch","mpl", and"posterior"under brms. The default is the Monte-Carlo interval, built by drawing parameter values from the fitted model’s uncertainty. Where a method does not apply to the design, the call aborts with a classed error. Confidence-interval methods compares them;?iccgives the per-method conditions. - Interval coverage was studied by simulation before release, and the limitations those studies found are documented rather than smoothed over. The Monte-Carlo default under-covers when the subject effects are strongly skewed or heavy-tailed, worst 0.6725 at chi-square(1) subject effects with a true ICC of 0.6, 50 subjects and 5 raters.
"burch"is no remedy there (worst 0.6655). -
"searle"and"burch"are the two classical closed forms, and the three grids below measure their widths: the smaller grid’s 16 cells and the larger grid’s 64 cells draw only the subject effects from the non-normal family. The third draws the residual from the same family as the subject effect. - Which of the two closed forms gives the narrower interval is conditional. On both grids that vary only the subject effect,
"burch"is the narrower of the two in 16 of 16 cells of the smaller grid and 59 of 64 cells of the larger grid. - That width margin holds much the same up to a true ICC of 0.3 rather than shrinking as the true ICC rises (on the larger grid; the smaller grid’s margin does shrink across its levels). The
"burch"width advantage then collapses to near parity at a true ICC of 0.6, on the one grid reaching that value. It also shrinks steadily as the subject count grows, measured at 5 raters. - What
"burch"does against"searle"also depends on the residual, and the three grids measure that: the two grids that vary only the subject effect put it narrower nearly everywhere, while the third, which draws the residual from the same family as the subject effect, puts it wider at every symmetric heavy-tailed family measured (a median width ratio of 1.2963 at t(5) with 100 subjects) and narrower at every lighter-tailed one, the normal included. - The
ci_method = "mpl"documentation states the interpolation evidence behind off-node subject counts. The correction constant is calibrated at subject-count nodes and linearly interpolated between them. The interpolated path is coverage-validated at each supported confidence level, at the default 0.95 by three off-node cells, each clearing its pre-registered coverage floor. The documentation also says what that validation does not establish (interpolated values are validated at a handful of geometries, not calibrated, and the interval’s asymmetry direction is not uniform across rater counts). Every universal or negative claim this documentation makes about the validated cells is settled mechanically, in CI, against the committed coverage fixture (data-raw/check-mpl-doc-claims.py).
