Supplemental Materials

Evaluating Open-Weight Large Language Models for Structured Depression Assessment from Clinical Interviews

This site hosts the supplemental analysis reports and source code for:

Girard, J. M., Kebe, G. Y., Morency, L.-P., De la Torre, F., Liebenthal, E., & Baker, J. T. (in press). Evaluating open-weight large language models for structured depression assessment from clinical interviews. Journal of Psychopathology and Clinical Science.

A preprint of the article is available on PsyArXiv.

Archived Version

Frozen, citable copies of these materials are archived on Zenodo. The version cited in the article is v1.1.0 (doi:10.5281/zenodo.23029765), which adds the model predictions to the analysis code and reports archived in v1.0.0.

DOI badge for version 1.1.0

Analysis Reports

Each report is a rendered Quarto document containing the complete analysis code and output. The corresponding source files (.qmd) are available in the source folder.

Report Description
Performance Analyses Benchmarking of 25 open-weight LLMs on MADRS item and total score prediction; direct vs. indirect scoring comparison; error characterization
Fairness Analyses Psychometric fairness audit (bias, sensitivity, precision) across sex, age, race, education, primary diagnosis, and interviewer
Ablation Analyses Prompt ablation study isolating the contributions of descriptive and demonstrative cues
Ensemble Analyses Comparison of cross-model ensemble aggregation to the best single model
Sample Demographics Participant and session descriptive statistics (Table 1)

Data Availability

The model predictions analyzed in this work are openly available (CC BY 4.0). The codebook describes every column and how to link the files to the NDA records.

File Contents Size
predictions.csv Predictions from all 25 models (three generations each) for the ten MADRS items and total score, plus the prompt ablations 7.7 MB
sessions.csv Interviewer code for each of the 541 sessions 7.5 KB
fewshot_exemplars.csv Sessions used as worked examples in the prompts, by target 1.2 KB

The complete v1.1.0 archive (data, code, and reports) can also be downloaded as a single file from Zenodo: Download all (ZIP, 5 MB)

The interview transcripts, human MADRS ratings, and participant characteristics needed to reproduce the analyses are available to qualified researchers through the NIMH Data Archive (NDA) in collection 3860 under an NDA Data Use Certification (how to request access). The public predictions link to these records through their NDA subject identifiers; the codebook lists the NDA structures and fields used. The interview recordings cannot be shared because they contain identifiable protected health information from psychiatric inpatients.

Contact

Jeffrey M. Girard, Department of Psychology, University of Kansas — jmgirard@ku.edu