# Lesson 7: Measurement and Psychometrics

*Companion-podcast transcript • Sarah & Kiffer*  
*Office Hours episode to listen to after working through the lesson*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This episode goes with Lesson seven, Measurement and Psychometrics.

**Sarah:** It's meant for after you've worked through the lesson. We'll bring in some perspective and some critique, spend time on the questions students tend to find thorny about scales and scores, and work through some extra examples, a few of them harder than the ones on the lesson page.

**Kiffer:** There are a few places where we ask you to work something out before we give the answer. When we do, you'll hear a few seconds of quiet. Pause the audio if you'd like more time.

**Sarah:** Here are the questions. If a scale has a high alpha, does that mean its items measure one thing? How much does an unreliable score hide, and can you correct for it? Does a scale measure everyone equally well? And we'll finish with a scoring rule that Kiffer and I don't entirely agree about, which is whether "more or less" should count as a lonely answer.

**Kiffer:** Let's start with alpha.

**Sarah:** Question one. For a long time I read a high alpha as a sign that the items all measure the same thing. That seems to be how a lot of published papers use it, too.

**Kiffer:** It's a very common reading, and numbers show why it doesn't hold. Here's a made-up scale with ten items. Five of them ask about loneliness, and the other five ask about something unrelated, say how much a person enjoys cooking. Within each set of five, every pair of items correlates 0.5. Between the two sets, every correlation is zero. So this scale measures two separate things.

**Sarah:** And we can get alpha from the standardized formula, because it only needs the number of items and the average correlation between pairs of items.

**Kiffer:** Right. So here's the question for everyone. Would this ten-item scale reach the usual guide of 0.70 for alpha? Take a few seconds.

*(Pause)*

**Kiffer:** Let's count the pairs. Ten items give forty-five pairs, which is ten times nine, divided by two. Each set of five has ten pairs, so twenty pairs correlate 0.5, and the other twenty-five pairs correlate zero. The average correlation is twenty times 0.5, which is ten, divided by forty-five. That's two ninths, or about 0.22.

**Sarah:** Then the top of the formula is ten times two ninths, which is about 2.22. The bottom is one plus nine times two ninths, which is exactly three. So alpha is 2.22 divided by three, or about 0.74.

**Kiffer:** So it clears the guide, for a scale built from two constructs that have nothing to do with each other. If the two sets correlated a little, say 0.2, the average correlation would rise to one third, and alpha would be about 0.83.

**Sarah:** So the length of the scale hides the split.

**Kiffer:** Yes. There's one clue in the numbers, though. Each set of five on its own has an alpha of about 0.83, which is higher than the 0.74 for all ten items together. When the alpha for a whole scale is lower than the alpha for one of its parts, the parts may not belong together. The six De Jong Gierveld items show the same pattern, with 0.59 for all six against 0.75 for the three social items.

**Sarah:** Although when the two sets correlate 0.2, the whole scale and each set of five both come out at about 0.83, so that clue disappears. The way to check is a factor analysis.

**Kiffer:** Right, so the clue is only a first sign. My advice is to read alpha as a summary of how strongly the items move together on average, and to test the structure with factor analysis. It's also worth knowing that Cronbach himself came to have doubts. In 1997 he began planning notes on how his views had changed, and they were published in 2004 with editorial help from Richard Shavelson. In them he doubted that alpha was the best way of judging the reliability of an instrument.

**Sarah:** So when I read a Methods section that says "alpha was 0.85, showing that the items measure a single construct", what should I look for?

**Kiffer:** Look for evidence from a factor analysis, ideally a confirmatory one, or at least the alphas and correlations for any subscales. If the paper gives only alpha, the claim that the items measure one construct hasn't been tested.

**Sarah:** Question two. How much does an unreliable score hide? This one goes straight back to Amira's reviewer, because the emotional subscale has an alpha of only 0.51.

**Kiffer:** Epidemiology has a long history with this problem, under the name regression dilution. A single blood pressure reading varies from day to day, so it's an unreliable measure of a person's usual blood pressure. In 1990, MacMahon and colleagues combined nine prospective studies with about 420,000 people. After they corrected for that random variation, the associations of diastolic blood pressure with stroke and coronary heart disease were about sixty percent greater than in earlier uncorrected analyses.

**Sarah:** So the arithmetic can run in the other direction. If error shrinks a correlation by a known amount, you can estimate how big the correlation would be without the error.

**Kiffer:** That's the correction for attenuation. You take the attenuation formula from the lesson and solve it for the true correlation. The corrected correlation is the observed correlation divided by the square root of the product of the two reliabilities.

**Sarah:** Can I try it on Amira's data? The social subscale correlates minus 0.48 with perceived social support, and the emotional subscale correlates only minus 0.26. A skeptical reader might say the emotional subscale looks less related to support simply because it's measured so poorly.

**Kiffer:** That's a fair challenge, and this is the right tool for it. Use the alphas as the reliabilities, 0.75 for the social subscale and 0.51 for the emotional subscale. We also need a reliability for the support score, so let's make one up and say 0.90.

**Sarah:** So here's the question for everyone. What's the corrected correlation between the emotional subscale and social support? Take a few seconds.

*(Pause)*

**Sarah:** Okay. The product of the reliabilities is 0.51 times 0.90, which is 0.459. The square root of that is about 0.68. Then minus 0.26 times 0.68 gives about minus 0.18.

**Kiffer:** Before we go on, does that answer make sense?

**Sarah:** Hmm. No, it doesn't. Measurement error pulls correlations toward zero, so correcting for it should move the correlation away from zero. I've made it smaller. I multiplied when I should have divided.

**Kiffer:** Exactly. A quick check of the direction catches that kind of mistake.

**Sarah:** So it's minus 0.26 divided by 0.68, which is about minus 0.38. For the social subscale, the product is 0.75 times 0.90, or 0.675. The square root is about 0.822, and minus 0.48 divided by 0.822 is about minus 0.58.

**Kiffer:** So after correction, it's about minus 0.58 for the social subscale and minus 0.38 for the emotional subscale. What would you tell the skeptical reader?

**Sarah:** That the gap is still there. It was 0.22 before the correction and it's about 0.20 after, so the difference between the subscales can't be put down to the emotional subscale's poor reliability. That fits Weiss's distinction, in which social loneliness is the kind that's tied to the wider network.

**Kiffer:** That's the right reading. Now for the trap. The correction is only as good as the reliabilities you feed into it. Alpha underestimates the emotional subscale's reliability, because its loadings are so unequal. With omega, which is 0.57, the corrected value is about minus 0.36. And the alphas were computed from the three-point item codes, while the subscale scores use the zero-or-one rule, so both reliabilities are approximate.

**Sarah:** And if you underestimate a reliability, the correction comes out too big.

**Kiffer:** Yes. With small samples or underestimated reliabilities, a corrected correlation can even come out above one, which is a clear warning sign. The correction also assumes that the errors in the two measures are unrelated. Both scores here come from the same online questionnaire, answered by the same person at the same sitting, so some of the error may be shared. My advice is to report the observed correlations as the main result and the corrected ones as a sensitivity analysis, naming the reliabilities you used.

**Sarah:** Does it matter which variable is measured badly? Amira might use emotional loneliness as the exposure in a regression with depressive symptoms as the outcome.

**Kiffer:** For a correlation, error in either variable weakens it. For a regression slope, it matters a great deal. Random error in the exposure flattens the slope by a factor equal to the exposure's reliability. Random error in the outcome leaves the slope unbiased and widens its confidence interval.

**Sarah:** So with a reliability of 0.51, a made-up true slope of 0.8 would show up as 0.8 times 0.51, which is about 0.41. That's roughly half the true slope.

**Kiffer:** Right. The same logic applies to confounders. If you adjust for a confounder that's measured with a reliability of about one half, you remove only part of the confounding it causes, and the rest stays in your estimate. Epidemiologists call that residual confounding, and it's one more reason to report the reliability of every scale that goes into a model.

**Sarah:** Question three. Does a scale measure everyone equally well? Classical test theory gives one reliability for the whole sample, so it seems to say yes.

**Kiffer:** It does seem to. A reliability of 0.75 leads to one standard error of measurement for everyone in the sample. Item response theory shows that precision changes along the trait. Let's take three made-up items, each with a discrimination of two and a difficulty of zero. That's close to the three social items under the published zero-or-one scoring, whose difficulties sit a little below zero.

**Sarah:** And each item's information is the discrimination squared, times the probability of the keyed answer, times one minus that probability. So at a theta of zero, where the probability is one half, each item gives four times a half times a half, which is one. Three items give a test information of three, and the standard error is one divided by the square root of three, which is about 0.58.

**Kiffer:** Right. Now take someone at a theta of two, which is two standard deviations above average in social loneliness. For that person, the probability of the keyed answer on each item is about 0.982. Each item gives four times 0.982 times 0.018, which is about 0.07. Three items give about 0.21, and the standard error is one divided by the square root of 0.21, which is about 2.2.

**Sarah:** On a trait where most people fall between minus three and three, a standard error of 2.2 tells you very little about that person.

**Kiffer:** That's right. These items separate people around the middle, and nearly everyone well above that level scores a point on all three. So here's a question for listeners. Suppose you could add one new item with the same discrimination, to measure the most isolated people better. Where should its difficulty be? Take a few seconds.

*(Pause)*

**Sarah:** Near a theta of two, where we just saw how poor the precision is. That would be an item that only quite isolated people endorse, something like "I have no one at all I can rely on".

**Kiffer:** Yes. An item gives its most information at its own difficulty, so an item with a difficulty of two adds an information of one at that level. The total rises from about 0.21 to about 1.21, and the standard error falls from about 2.2 to about 0.91. It barely changes the precision in the middle of the trait.

**Sarah:** So the single reliability of 0.75 for the social subscale summarizes a precision that's good for some people and poor for others.

**Kiffer:** Yes. Classical reliability gives one summary for the whole sample, and item response theory breaks it down by level of the trait.

**Sarah:** Which is the idea behind adaptive testing. If you know roughly where a person sits, you can give them the items that measure well at that level.

**Kiffer:** It is. The PROMIS measures used in health research are built with item response theory, and their computer adaptive versions pick each new item from a large bank, based on the person's earlier answers. A standard adaptive test gives between four and twelve items.

**Sarah:** This connects back to the loneliness scale in a way I didn't expect. Under the zero-or-one rule, the social items all have difficulties a little below zero. In the graded response model, which keeps all three answers, the second thresholds sit between about 1.0 and 1.3.

**Kiffer:** That second threshold separates "more or less" from the lonely answer. The zero-or-one rule merges those two answers, so it discards the part of each item that measures the lonelier end of the trait. That fits the distribution of total scores, too, since nearly half the people with complete answers score five or six out of six.

**Sarah:** Which brings us to the scoring rule, and the part where we disagree.

**Kiffer:** Question four, then. Should "more or less" count as a lonely answer?

**Sarah:** Let me set it up with a hypothetical respondent. She answers "more or less" to all six items. She more or less feels a sense of emptiness, she more or less has people she can rely on, and so on. What total score does she get under the published rule? Take a few seconds.

*(Pause)*

**Kiffer:** She gets six, the maximum. Each "more or less" scores one point, so she has the same score as someone who gives the lonely answer to every item.

**Sarah:** And that's my concern. Someone who is ambivalent about every item is scored as lonely as a person can be. If you add the three-point codes, she scores twelve on a range from six to eighteen, and the person with every lonely answer scores eighteen. The published rule erases that difference.

**Kiffer:** The scale's authors made that choice deliberately. Their manual states it as an assumption, drawn from their earlier work, that the middle answer indicates loneliness. They also prefer the zero-or-one scores because those make results comparable with earlier studies.

**Sarah:** I accept the point about comparability. But look at what happens at the cut-point. The manual classifies scores of two to six as lonely, and in the 2021 wave of the Canadian Social Connection Survey, that's about ninety percent of the people who answered all six items. In 2021, Statistics Canada's Canadian Social Survey found that thirteen percent of people aged fifteen and older always or often felt lonely.

**Kiffer:** That comparison mixes several things, though. The Canadian Social Connection Survey recruited volunteers online, the Statistics Canada figure comes from a different question, and "always or often" is a much higher bar than a score of two. So the gap says little about whether the De Jong Gierveld scale ranks people well. Where I agree with you is that a cut-point builds a judgement into the prevalence figure that gets reported, and the manual itself calls its cut-points tentative.

**Sarah:** I'd go further. The cut-point is a choice, and so is the rule underneath it. Given what the item response curves showed, I'd make the three-point score the main measure, because it keeps the distinction between "more or less" and the lonely answer.

**Kiffer:** I'd keep the published rule as the main score. I find the authors' assumption reasonable for the positively worded items. If the most a person can say to "There are plenty of people I can rely on when I have problems" is "more or less", that hesitation probably tells you something about their network. And a main score that few other studies use makes your results hard to compare with earlier work.

**Sarah:** That's fair for comparisons between studies. My worry is the prevalence figure, because that's the number that ends up in news stories and in decisions about funding.

**Kiffer:** Then I think we agree on the practical rule. Score the scale with the published rule so your results can be compared. Report the distribution or the mean alongside any prevalence, and when you do report a prevalence, give the cut-point and the question wording. Then check your main results with the three-point score as a sensitivity analysis.

**Sarah:** And never set a prevalence from one instrument beside a prevalence from another as if they measured the same thing.

**Kiffer:** I agree with that completely. Let's pull it together with three things to take away.

**Sarah:** First, alpha summarizes how strongly items move together on average. It rises with the number of items, so a long scale built from two unrelated sets of items can still pass the usual guide, and the structure has to be tested with factor analysis.

**Kiffer:** Second, unreliable scores weaken associations, and the correction for attenuation estimates by how much. The corrected value depends on the reliabilities you use, so report it as a sensitivity analysis beside the observed result.

**Sarah:** And third, precision and meaning both depend on how a scale is scored. Item response theory shows where on the trait a scale measures well, and a scoring rule or a cut-point carries a judgement that the Measures paragraph should state.

**Kiffer:** If you'd like more practice, redo the correction for attenuation from question two with omega as the emotional subscale's reliability, and then with a support reliability of 0.80. Each time, check that the corrected value moves away from zero, and see whether the gap between the subscales survives.

**Sarah:** Next time, it's Lesson eight, Mediation, Moderation and Path Analysis, where we ask what role a third variable plays in an association, and the factor model from this lesson becomes part of a structural equation model.

**Kiffer:** Take care, everyone.

**Sarah:** See you in Lesson eight.
