The invisible thing, measured three ways

multiverse
psychometrics
Sum scores, item response theory, and network psychometrics are three theories of what a questionnaire just did.
Author

Constantin Yves Plessen

Published

August 24, 2026

Before we even think about researcher degrees of freedom, all those paths we could walk down, smart people already went hiking when they created the instruments we need to use in our research project. Say 9 questions on depressive symptoms like sleep, mood, appetite, with 4 possible answers, ranging from 0 = no problems to 3 = very severe problems. We then count those answers and sum them up with a possible range of sum scores from 0 to 27.1 In this approach, one point of sleeplessness is equal to one point of depressed mood, which is equal to one point of suicidal ideation. Which many researchers consider weird for many different reasons. For instance, many people have regularly increased or decreased appetite (I have that multiple times daily), or slept bad (I have a newborn next to me currently familiarizing himself with the power of his vocal cords, so this question needs context), and unfortunately still too many people suffer from suicidal ideation, or depressed mood and sadness. But they are not the same thing or interchangeable.2 Even though all of this is problematic, tests built like this in the tradition of classical test theory are very useful, pragmatic, and luckily correlate insanely well with much more advanced psychometric techniques like item response theory and network theory.

Item response theory would not count each question equally, but would rather consider some questions to measure high levels of depression while other questions would measure lower levels. IRT has clear strengths when measuring IQ and educational tests, where questions can be easier or more difficult in an obvious way: 2+3 tends to be easier for most than 3842^33292* pi. If you can’t answer the first, you are in trouble, if you can answer the second, you might also be in trouble, but for a very different reason. For mental disorders I doubt that such a clear seperation is possible, yet there are still many interesting things you can do once you applied IRT for instrument development: you can investigate whether the test measures the same construct across cultures or groups, does the test perform similarly in men and women? When asking for knee functioning, the item “can you go to the toilet without problems” measures different things, as women tend sit down and men stand up most of the time. When you ask “can you dance passionately” an Argentinian might have a different picture in mind than a Bavarian (also vigerous exercise, but in my mind, very different movement patterns).3 IRT can test for that.4 You can create computer adaptive tests that only give the next best item based on the estimated ability a person has and reduce a 100 item questionnaire to ~10, reducing the burden on the person substantially and potentially producing better data, as the person won’t click the middle answer category due to boredom and spite. The only downside is: it costs an ungodly amount of money and time to create the required collection of items, calibrating them, validating them, translating them.

Network psychometrics would disagree with both approaches on a philosophical level: there is no latent construct like depression lurking in the shadows and pulling its statistical strings: there are symptoms that might reinforce and cause each other: can’t sleep? Less energy. Less energy? Lower mood. Lower mood? Less urge to call some friends over. Less friends over? More loneliness. More loneliness? More worrying. Worrying? Can’t sleep. Can’t sleep? …5

In this theory, depression would be the configuration of that network of symptoms. The overall sum score would not be interesting, as the many connections are the core thing: what changes, why, how. Where should interventions start? Treat sleep or social isolation first? In this framework, this is a testable empirical question. Yet again, insanely expensive as thousands of participants are needed to investigate, validate, and improve those network models.

Those new fancy models offer amazing opportunities to improve our knowledge of how therapies could be personalized, how measures could measure more fairly and efficiently, or even how mental disorders arise. However, many stakeholders see only little benefit over old school sum scores.6 After doing some fancy IRT modeling I was so proud of I correlated the resulting T-score to the sum score and was disappointed that all that work led to a correlation of >.95.7 So I can understand the critique that IRT might be a bit over the top for many applications. Symptom networks are quite difficult to interpret and many reviewers are not familiar with them. So it seems there is this chasm between the old, pragmatic, oftentimes spot on, difficult to validate, and theoretically “wrong” sum score, the insanely expensive and powerful IRT models, and the very interesting and difficult to fund network psychometrics. All three approaches are very reasonable to some, indefensible to others, and for outsiders who just want to use a good questionnaire to investigate their research question not really answerable. All those theories seem to be reasonable, yet to many researchers it us unknown wether choosing a differnt framwork changes only an arbitrary method, or the question, or even the estimand, the very thing we are trying to measure.8

AI disclaimer: Written by me without LLM assistance.

Layer 1 — Psychometrics. The annotated literature behind this essay — measurement, sum scores, IRT, and network models — lives in the living collection.

Footnotes

  1. This describes the PHQ-9: Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2001). The PHQ-9: Validity of a brief depression severity measure. Journal of General Internal Medicine, 16(9), 606–613. https://doi.org/10.1046/j.1525-1497.2001.016009606.x↩︎

  2. The argument at length: Fried, E. I., & Nesse, R. M. (2015). Depression sum-scores don’t add up: Why analyzing specific depression symptoms is essential. BMC Medicine, 13, 72. https://doi.org/10.1186/s12916-015-0325-4↩︎

  3. The two movement patterns in question: ↩︎

  4. This is called differential item functioning (DIF). We have mapped how much its detection itself depends on analytic choices — across the full space of defensible DIF specifications — in PROMIS physical function items: Plessen, C. Y., Fischer, F., Hartmann, C., et al. (2024). Differential item functioning between English, German, and Spanish PROMIS physical function ceiling items. Quality of Life Research, 34(5). https://doi.org/10.1007/s11136-024-03866-y↩︎

  5. Borsboom, D. (2017). A network theory of mental disorders. World Psychiatry, 16(1), 5–13. https://doi.org/10.1002/wps.20375; Borsboom, D., & Cramer, A. O. J. (2013). Network analysis: An integrative approach to the structure of psychopathology. Annual Review of Clinical Psychology, 9, 91–121. https://doi.org/10.1146/annurev-clinpsy-050212-185608↩︎

  6. Byrne, et al. (2026). ‘We’re not going to start lifting stones now…’: Stakeholder perspectives on the role of psychometric methods in outcome measurement. British Journal of Clinical Psychology. https://doi.org/10.1111/bjc.70067↩︎

  7. We found the same pattern in real trial data — re-scoring an RCT’s outcomes with IRT barely moved the conclusions: Harrison, C. J., Plessen, C. Y., Liegl, G., et al. (2023). A psychometric sensitivity analysis of the TOPKAT trial. Journal of Clinical Epidemiology, 158, 62–69. https://doi.org/10.1016/j.jclinepi.2023.03.013↩︎

  8. This question — when does a fork change the method, and when does it change the question being asked — is the live debate: Short, C. A., et al. (2026). Multicurious: A multidisciplinary guide to multiverse analysis. Advances in Methods and Practices in Psychological Science. https://doi.org/10.1177/25152459261434881; Auspurg, K., & Brüderl, J. (2021). Has the credibility of the social sciences been credibly destroyed? Socius, 7. https://doi.org/10.1177/23780231211024421↩︎