The same clinician, twice: a paired look at Study 1

Self-paced bonus activity · what a within-subjects design buys you, and what it costs

Welcome to a self-paced bonus activity. In the M09 lab, Group A ran a two-sample Welch’s t-test on Study 1 and found that clinicians shown standard-error bars estimated a much higher probability of superiority than clinicians shown standard-deviation bars.

That analysis used one estimate per clinician. But Study 1 was a within-subjects design — every clinician eventually saw both figures and gave two estimates. Here you’ll run the paired analysis the lab set aside, and find out two things that are genuinely surprising:

  1. Pairing does buy precision here — but only through one of the two routes it normally works by. The other route is missing entirely, and you can see why.
  2. The size of the within-person effect depends on presentation order, which you can test directly. That is a display × sequence interaction, and it is why the first-estimate comparison is the cleaner one.

Everything happens in the sandbox chunks on this page — nothing here goes into your lab notebook. The blood-pressure subset (paired_bp) is already loaded, and tidyverse, infer, and effectsize are attached.


The data, in wide form

The lab’s file kept each clinician’s first answer. This one keeps both, one row per clinician:

Two columns worth reading carefully. psup_se is the estimate a clinician gave for the SE figure and psup_sd the estimate for the SD figure — whenever in the session each appeared. The separate first_condition column records which came first. So for a clinician who saw SDs first, psup_sd is their first answer and psup_se their second.

Your task: compute the within-person difference and summarize it. The difference we want is psup_se - psup_sd, so that a positive value means this clinician gave a higher estimate for the SE figure.

Positive should mean “higher for SE,” so psup_se goes first.

On average, clinicians gave estimates more than 20 percentage points higher for the SE figure than for the SD figure. Same clinician, same underlying study result, two different uncertainty displays. But because one display necessarily came second, this raw within-person contrast can also carry an order or carryover effect — a possibility we’ll test below.

How lopsided is it?

A mean can hide a lot. Count the directions:

Out of 75 clinicians, the overwhelming majority moved the same way and only a handful moved against the effect. That is a stronger statement than the mean alone can make — it shows the positive average is not produced by a small minority moving a long way. Note what it does not settle: the magnitude of the mean difference can still be influenced by how far individual clinicians moved, and counting directions deliberately ignores distance.

The paired test

Here is the key move: a paired t-test is a one-sample t-test on the difference scores, tested against zero — the same one-sample t the Module’s on-ramp walked through, just applied to a column of differences. That also tells you which distribution the model considerations are about: the difference scores, not the two raw outcomes separately. With 75 differences the mean is fairly robust, though a few extreme differences could still dominate it; the scatterplot further down shows each one as a point’s vertical distance from the dashed line.

Your task: two blanks — the response is the diff column you just built, and the null value is 0.

Base R’s t.test() will do the subtraction for you if you hand it both columns and set paired = TRUE — same answer, less bookkeeping:

Our estimate, tobs, and df match the authors’ exactly

The paper’s main text mentions the within-subjects result only in passing (SI Appendix, Fig. S3), but the authors publish their full analysis code and its rendered output in the repository. Their within-subjects table reports:

Scenario Mean difference tobs df
Blood pressure 22.20000 9.419171 74
COVID-19 18.35795 9.083092 87

Your blood-pressure numbers above should match to every printed digit. One difference to expect: the authors ran their test one-sided (alternative = "greater"), having predicted the direction in advance, so their p-value and confidence bound differ from the two-sided ones you just computed. The estimate, the statistic, and the df are identical — only the tail area changes. That is the α and one-sided-versus-two-sided discussion from Part 1 of the Module, showing up in a real paper.

They also report a sign test as a distribution-free check (blood pressure: \(p = 1.19 \times 10^{-13}\)). It uses only the directions of the nonzero differences and discards their magnitudes, so it needs no Normal or even symmetric difference distribution. It is not assumption-free, though: it still requires independent pairs, a stated rule for handling ties (clinicians who gave identical answers), and, under its null, an equal chance of a positive or negative difference.

A common standardized effect for a paired design is Cohen’s \(d_z\) — the mean difference divided by the SD of the difference scores. It is not the only defensible standardizer for repeated measures, which is why effectsize points you toward repeated_measures_d() for the alternatives:

Surprise 1 — the usual covariance advantage is absent

Here is where this dataset stops behaving like a textbook example — though not in the way you might expect.

With a fixed number of participants, a repeated-measures design can improve precision in two separate ways.

  1. Participant efficiency — every participant contributes under both conditions. The same 75 clinicians give you 75 observations of each display, 150 in all. The between-groups comparison, by contrast, splits those same 75 people into two independent groups — 41 and 34 — so each display is observed only about half as often.
  2. Covariance efficiency — positive within-person correlation cancels stable individual differences. Whatever makes a given clinician generally optimistic sits in both of their answers and subtracts away. This is the effect textbook examples usually showcase, because they usually feature strongly correlated repeated measures.

Those are independent. Route 1 is about how many observations you collect from the people you recruited — compare designs with the same number of observations per condition and it disappears. Route 2, the statistical advantage of pairing itself, depends on the correlation between the two measurements. Check the correlation here:

The correlation is essentially zero, so there is little to no covariance advantage in this sample: knowing a clinician’s SE answer provides almost no linear prediction of their SD answer.

Notice that sd_diff comes out larger than either raw SD. That is not evidence of anything going wrong — it is arithmetic. Since

\[ \text{Var}(X - Y) = \text{Var}(X) + \text{Var}(Y) - 2\,\text{Cov}(X, Y) \]

a correlation near zero makes the covariance term very small, so the variance of the difference is close to the sum of the two variances — which exceeds either one. Comparing sd_diff to a raw SD therefore tells you nothing about precision. Precision is about the standard error of an estimator, not the SD of a score — so compare the standard errors directly.

Almost every point sits above the diagonal — the effect is overwhelming — but the cloud has almost no linear upward tilt, consistent with the near-zero correlation you just computed.

Now compare the two standard errors

This is the comparison that actually settles whether pairing bought precision. Put the paired estimator’s SE next to the between-groups (Welch) SE for the same 75 clinicians:

Pairing did buy precision — through one route, not two

The paired standard error is about 71% of the between-groups standard error. That is a real precision gain: the paired 95% interval is roughly nine percentage points wide against thirteen for the between-groups comparison, which contributes to the paired tobs (9.42) being larger than the lab’s between-groups tobs (6.83).

One caution. Once we discover below that the contrast varies by presentation order, these two analyses are not estimating exactly the same quantity: the paired estimate averages the SE-versus-SD contrast across both sequences, while the between-groups estimate uses first presentations only. So this comparison is a useful illustration of the efficiency gained by collecting both conditions from every clinician — not a pure head-to-head of two estimators of the identical estimand.

So the headline is not that pairing failed. It is that pairing delivered route 1 and only route 1:

Operating here? Why
Route 1 · participant efficiency — everyone contributes under both conditions ✅ fully each display observed 75 times, instead of 41 and 34
Route 2 · covariance efficiency — positive correlation cancels individual differences ❌ essentially none r ≈ -.05, so essentially nothing is being cancelled through positive covariance

The lesson worth keeping: “paired designs are more precise” is really two claims, and they can come apart. With the same participant budget, this repeated-measures design gained precision through participant efficiency, and essentially none through positive covariance. Textbook examples usually feature strongly correlated repeated measures, so both routes fire at once and the two get conflated. Here you can watch them separate.

One thing this is not: a rule for choosing your test. The design decides that — these are two measurements on the same clinicians, so the paired analysis is the one that matches the data structure. What the correlation explains is how much precision you ended up with, not which test was appropriate. Prior evidence about within-person correlation is useful when you’re planning a study and want to anticipate the payoff from a repeated-measures design.

Surprise 2 — the effect depends on presentation order

Now the reason the first-estimate comparison is the cleaner one.

Every clinician answered about one figure, then about the other. If the sequence they happened to get changed how they answered, then the within-person difference is not a pure display effect — it also carries whatever the sequence did.

Order was randomized, which gives us a way to check. If sequence made no difference, the within-person contrast should be the same size for clinicians who saw SEs first as for those who saw SDs first.

Your task: one blank — group by first_condition.

An exploratory follow-up. Those two means look surprisingly far apart. Because we noticed this pattern in the data rather than specifying the test in advance, treat what follows as exploratory — the M08 rule about deciding before you look applies to follow-ups too:

The within-person effect is substantially larger for clinicians who happened to see the SE figure first, and the difference between those two within-person effects is itself statistically significant.

What this provides evidence for, stated carefully: the size of the within-person display contrast differs by presentation sequence. In model terms, that is a display × sequence interaction, and it means a single pooled paired estimate is averaging over two genuinely different quantities.

What it does not establish is the mechanism. Any of these would produce the same signature, and this test cannot tell them apart:

  • Anchoring — the first answer constrains the second.
  • Carryover — seeing one figure format changes how the next one is read.
  • A period effect — second answers differ systematically from first answers for reasons unrelated to format (fatigue, practice, a shifted sense of the scale).
  • Ceiling effects — the scale tops out at 100, so clinicians who started high have less room to move.

Distinguishing those explanations would require additional modeling and, for some of them, additional design information or data; this two-period contrast cannot tell them apart. It is enough — and more honest — to say the contrast varies by sequence and name the candidates.

What this explains about the paper

The first-estimate comparison sidesteps all of this: it compares randomized groups before either clinician has seen the alternative figure, so there is no sequence for order to act through. That is what makes it clean.

Zhang et al. do lead with the first-estimate comparison, and report the within-subjects analysis only in their SI Appendix, noting it shows “a similar pattern.” Our order result is consistent with that choice — but the paper does not state this as its reason, so treat the connection as a reasonable interpretation rather than something the authors said.

Both analyses are defensible, and they answer different questions. The paired one carries a caveat the between-groups one does not:

Between-subjects (the lab) Paired (this page)
Question Do two groups of clinicians differ? Does the same clinician answer differently?
Uses first estimates only both estimates
Clean? yes — no prior figure could have influenced the answer the contrast varies by sequence (shown above)
The paper headline result SI Appendix

The transferable habit: when a design gives you more than one way to analyze it, the right choice is rarely “whichever gives the bigger tobs.” It is whichever question you can answer most cleanly — and saying out loud what the other analysis showed, rather than quietly reporting only one.

Try it on the other scenario

Study 1 randomized clinicians to a blood-pressure or a COVID-19 trial. Everything above used the blood-pressure arm. Rerun it on the other one and see whether the same two surprises hold:

The same paired pattern appears in the COVID-19 arm. Whether the sequence effect does is worth checking for yourself — and is a good reminder that a pattern found in one arm of a study is a hypothesis about the other, not a finding.


Back to the M09 lab → return to the lab