Conducting Common Hypothesis Tests

Pre-Study · Module 9 · Fri Oct 9

Welcome to the M09 pre-study. M08 taught you one testing logic, and the M09 Module showed that same logic at work across different research designs. This pre-study is where you put it into practice: each video walks through one piece step by step, and the activities after it have you run those same steps yourself in R.

How this page is organized

Five videos, one idea each:

  1. Every test statistic does the same job: it measures how far the data depart from what \(H_0\) predicts, relative to the departures the null model ordinarily produces. \(t_\text{obs}\) is your result; \(t^*\) is the cutoff you compare it with.
  2. A two-sample t-test compares a difference with its uncertainty: the observed gap between two groups, relative to ordinary sampling variation.
  3. Effect size answers a different question than a p-value: p is about compatibility with the null model; d is about how big the gap is.
  4. ANOVA compares between-group signal with within-group noise: when the null is true, F averages about 1, and a significant F says some means differ, not which.
  5. Chi-square compares observed counts with the counts independence predicts: and those expected counts also tell you whether the test can be trusted.

Some chunks between activities just run: they set up the data, check your answer, or extend it. Run each as you reach it, because the text below it depends on the result. Boxes marked Going further are extras, some open and some folded away; nothing later needs them.

No need to memorize formulas. For each test, aim to say what is being compared, what the null predicts, and what the result means. Each video section ends with a self-scoring Quick Check.

Plan on two or three sittings. This pre-study is longer than most, with five videos and eight activities. Each video section is a natural stopping point and works on its own when you come back to it. By the end, you’ll be able to choose, run, and read the three tests behind many published group comparisons in psychology.


Meet the data: the Hofman boulder-sliding study

Start with the paper. Every analysis on this page comes from one study: Hofman, Goldstein, and Hullman (2020), How visualizing inferential uncertainty can mislead readers about treatment effects in scientific results. Read it online or open the copy saved on this site (PDF). It’s worth reading in full: seeing how the authors framed their question, designed their experiments, and interpreted what they found is good practice for reading research, and for writing it.

A quick reminder. You met this study in the Module, and Video 2 walks through it again: people decided how much to pay to rent a “special boulder” in a fictional sliding game, after seeing either a confidence-interval (CI) chart or a prediction-interval (PI) chart of the same results. Their willingness to pay, in Ice Dollars, is the outcome in every analysis on this page.

Two experiments, and we use both. Experiment 1 crossed two factors: each participant saw one of two charts (CI or PI) with one of two captions, with the special boulder 4 m ahead. Experiment 2 crossed four displays (CI, PI, a rescaled CI, and HOPs, an animation that shows one simulated pair of outcomes at a time) with two effect sizes; in its large-effect condition the special boulder averages 116 m. Videos 2, 3, and 5 and their activities use Experiment 1 (blorg_exp1, described below). Video 4 and its activity use Experiment 2’s large-effect condition (blorg_exp2).

We analyze slices, not the full designs. Each test on this page compares groups on one factor, within a subset of the data: for example, Experiment 1’s matching-caption participants, or Experiment 2’s large-effect condition. That keeps every analysis to a single question. (The authors analyzed Experiment 1 the same way, with a separate t-test for each caption condition.) Models that handle both factors at once, including whether the effect of one factor depends on the other (an interaction), come in PSY 653.

blorg_exp1 · 1,743 observations · 6 variables · Hofman, Goldstein, & Hullman (2020), Experiment 1 · data/blorg_exp1.Rds

A 2 × 2 factorial: each Mechanical-Turk participant viewed one visualization comparing a “special boulder” (the treatment) and a “standard boulder” in the boulder-sliding cover story, then reported how much they’d pay (in Ice Dollars, 0–250) to rent the special boulder. The factors:

  • Visualization — 95% confidence interval (CI) error bars, or 95% prediction interval (PI) error bars.
  • Caption text — “matching” (caption describes only what the figure shows) or “extra info” (caption gives both 95% CI and 95% PI numerical bounds, regardless of figure).

That gives four conditions of roughly 400–500 participants each. Variables:

  • worker_id character — Anonymized Amazon Mechanical Turk worker identifier
  • condition.f factor — Four-level experimental condition crossing visualization format and caption text
  • interval_CI numeric — Visualization format shown to the participant
  • text_extra numeric — Whether the caption supplied information beyond the visualization
  • wtp_final numeric — Willingness to pay to rent the special boulder, in Ice Dollars
  • superiority numeric — Participant’s estimate of the probability that the special boulder out-slides the standard one

Full codebook for blorg_exp1 — values, levels, missingness, and how the file was prepared.

In every code chunk on this page, the dataset is available as blorg_exp1, with pre-built subsets mt (matching text, used throughout) and et (extra info, used in one Going further box). Experiment 2’s large-effect condition is pre-built as all_groups, Video 4’s data: two columns, graph_type (CI, CI rescaled, PI, HOPS) and wtp.


Video 1 — The machinery, and the two numbers you compare

What to listen for:

  • A difference means something only relative to the noise around it. The same 2-point drop in depression scores is a clear result when scores typically sit 1 point from their group’s average and a weak one when they sit 8 points away. The ruler is the null-world wobble: rerun the study with no real effect and the group difference still varies from run to run. Its typical size is the standard error, and \(t_\text{obs}\) counts how many of those wobbles your gap is (about 9.6 against 1.1 for the two samples).
  • Every test statistic in M09 does the same job: it measures how far the data depart from \(H_0\), relative to the departures the null model ordinarily produces. The arithmetic differs by test — \(t_\text{obs}\) is literally a gap over its standard error, while F and χ² aggregate departures in their own ways — but the gist is the same.
  • The four steps that turn any test statistic into a decision: calculate the test statistic, identify its reference distribution, determine the critical value, and apply the decision rule. The calculations change from test to test; the steps don’t.
  • The two symbols that get confused. \(t_\text{obs}\) comes from your sample. \(t^*\) is the cutoff the reference distribution supplies for your chosen \(\alpha\) and df — smaller df, heavier tails, a more extreme cutoff. Degrees of freedom usually come from counting: quantities you started with, minus constraints they must satisfy.

Going further — what df does to the cutoff. Video 1 showed three of these curves; here is the whole ladder as a loop, so you can watch \(t^*\) move. Same \(\alpha\), same rule — only the curve changes:

At df = 3 the critical value is 3.182; at df = 30 it is 2.042; by df = 500 it has settled at 1.965. At the same \(\alpha\), fewer degrees of freedom mean a larger critical value — and that isn’t a rule anyone imposed: it is the shape of the curve.

Quick Check

Answer each question — you’ll see green (correct) or pink (incorrect) feedback as you type.

1. What does tobs represent in a two-sample t-test?

2. Which of these is needed to compute t*, the critical value?

3. The same 2-point gap shows up in a tightly screened sample and in a heterogeneous sample of the same size. Why is it much stronger evidence in the tightly screened sample?

4. Suppose the new therapy truly has no effect, and the same randomized trial is run again and again. What happens to the difference between the two group means?

5. Video 1 lays out the steps every test in M09 follows to turn a test statistic into a decision. Which of these are among those steps? Choose Yes or No for each.

  • Calculate a test statistic
  • Choose α after looking at the results
  • Identify the reference distribution
  • Calculate the probability that H₀ is true
  • Determine the critical value
  • Apply the decision rule

6. As the degrees of freedom get smaller, what happens to t* (α = .05, two-sided)?

7. What job do degrees of freedom do in a hypothesis test?


Video 2 — The two-sample t, start to finish

What to listen for:

  • The design chooses the test. Four questions — what kind of outcome, how many groups, independent or paired, and which population quantity — point Hofman’s Experiment 1 to the two-sample t-test.
  • The hypotheses come first, and they are about populations. \(H_0\) says \(\mu_{CI} - \mu_{PI} = 0\). The alternative is two-sided, even though the authors expected CI viewers to pay more, and α = .05 is chosen before looking at any results, split 2.5% into each tail.
  • Name the estimand before computing anything. The estimand is the population gap, \(\mu_{CI} - \mu_{PI}\); the sample gap, \(\bar{x}_{CI} - \bar{x}_{PI}\) = $29.37, is our estimate of it.
  • Video 1’s four steps, on real data. The reference distribution is a t curve with df = \(n_{CI} + n_{PI} - 2\) = 903, with cutoffs at ±\(t^*\) ≈ 1.96. Then \(t_\text{obs}\), the gap over its standard error, comes out to 8.62 and lands far past a cutoff. The p-value comes from the same curve (the tail beyond the statistic’s size, abs(t_obs), from pt(), doubled for two sides), so the two readings agree: reject \(H_0\).
  • In practice, R runs the whole test. t_test() reproduces the video’s numbers when you add var.equal = TRUE; without it, R runs Welch’s version, which lets each group keep its own spread. Activities 2.2 and 2.3 put the two side by side. The output also gives a 95% confidence interval for the population gap, about $22.7 to $36.1, and because the interval and the test are built from the same estimate, standard error, and \(t^*\), 0 falls outside it exactly when the test rejects.

Activity 2.1 — Build the statistic ✍️

The matching-text subset is pre-built as mt (the rows where text_extra == 0), with one addition: graph_type, a factor labelling each person’s display CI or PI — the same name the Module uses, built from the dataset’s 0/1 interval_CI. Its descriptives are below. Read them carefully, because the task that follows uses them.

Your task: build the two pieces of the fraction. The gap is \(\bar{x}_{CI} - \bar{x}_{PI}\); the standard error of the difference is \(SE_\text{diff} = s_\text{pooled}\sqrt{1/n_{CI} + 1/n_{PI}}\), where \(s_\text{pooled}\) = $51.13 is the two groups’ SDs combined into one (it’s given here; the Going further box after this activity shows where it comes from).

Four blanks, all at the top: read each group’s mean and sample size off the table above (m_ci and n_ci come from the CI row). The lines below them then use those numbers to build the gap and its standard error, with nothing more to fill in.

The CI mean and sample size sit in the CI row of the table (columns M and n); the PI ones in the PI row. Type the numbers exactly as they appear, to two decimals.

You should get a gap of about $29.37 and a standard error of about $3.41. The inputs are rounded to the cent, but that barely matters here: built from the unrounded data, the gap, the standard error, and \(t_\text{obs}\) come out the same to two decimals. Typing numbers off a table is a good way to see how the fraction works; in a real analysis you’d let R pull them from the data instead, which is what t_test() does in Activity 2.3.

The pooled SD is one standard deviation built from both groups. Square each group’s SD to get its variance, weight each variance by its group’s degrees of freedom, average them, and take the square root:

\[ s_\text{pooled} = \sqrt{\frac{(n_{CI}-1)\,s_{CI}^2 + (n_{PI}-1)\,s_{PI}^2}{n_{CI}+n_{PI}-2}} \]

With the two SDs from the table above:

The weights are Video 1’s counting rule again. Each group spends one degree of freedom on its own mean, so the weights are 418 and 485, and the denominator is their total, 903. The PI group is larger, so its SD counts for a little more. That is why $51.13 lands between the two SDs, slightly nearer the PI group’s $49.38.

Hold onto the squared version, \(s_\text{pooled}^2\). In Video 4 it comes back as \(MS_\text{within}\): the same pooled spread of people around their own group’s mean, computed across four groups instead of two.

Think before you read on

Work this out in your head first, then click Show the answer below to check your thinking.

The gap is about $29.37. The standard error of that gap is under $4. But the within-group SDs are roughly $50 — far bigger than the gap itself. How can \(SE_\text{diff}\) be so much smaller than the SDs it was built from?

Because \(SE_\text{diff}\) is the standard error of a mean difference, not of an individual observation. A mean wobbles far less than a single person does, and the wobble shrinks as \(n\) grows — that is what the \(1/n\) terms do, and with \(n\) in the hundreds they shrink it hard:

\[ SE_\text{diff} = s_\text{pooled}\sqrt{\frac{1}{n_{CI}} + \frac{1}{n_{PI}}} \]

So \(SE_\text{diff}\) inherits the \(1/\sqrt{n}\) shrinkage from each mean. The statistic is large because the observed gap is large relative to its standard error — and with hundreds of people per group, that standard error is small. The raw gap alone cannot tell you how strong the evidence against \(H_0\) is.

Activity 2.2 — Cutoffs, \(t_\text{obs}\), and the p-value: standard and Welch ✍️

Video 2 drew the reference curve for this comparison, a t-distribution with df = 903, with its cutoffs at about ±1.96, then dropped our statistic onto it. Now build those numbers yourself. qt() and pt() are the t-family twins of M06’s qnorm() and pnorm(): qt() turns a probability into a cutoff, pt() turns a value into the area below it, and both take the degrees of freedom.

Your task: three blanks — the probability that gives the upper cutoff, the gap on top of \(t_\text{obs}\), and the value pt() needs to return one tail beyond the statistic.

qt() takes the area below a point and returns the point: the upper cutoff leaves .025 above it, so it has .975 below. \(t_\text{obs}\) is the gap from Activity 2.1, mean_diff, over its standard error. pt() normally returns the area below a value; lower.tail = FALSE flips it to the area above, the same switch you used with pnorm() in M06. Hand it abs(t_obs), the statistic’s size with its sign dropped, to get the tail beyond it, and doubling counts the matching tail on the other side.

You should get df = 903, cutoffs of ±1.963, \(t_\text{obs} \approx\) 8.62, and a p-value of about 3.1e-17. R writes very small numbers in scientific notation: e-17 means “move the decimal point 17 places to the left,” so this is a decimal point, sixteen zeros, then the digits. Both checks print TRUE: the statistic is past a cutoff, and the p-value is below α. And qnorm() gives 1.960: at df in the hundreds the t curve is essentially the Normal from M06, which is why Video 2’s cutoff looked like the familiar 1.96.

Notice what \(t^*\) did not need: the observed gap. It comes from α, the two-sided rule, and the df alone, and the df of 903 is Video 1’s counting rule: 905 observations, minus the two group means estimated from them. That is why the cutoff can be fixed before you see any results, while \(t_\text{obs}\) has to wait for them.

Why abs()? For a two-sided test, “at least as extreme” means at least as far from 0 in either direction, and abs(t_obs) is that distance, so the line is right whichever way you subtracted (PI − CI would give −8.62). One check worth keeping: a p-value can never be above 1; if you ever get one, you took the wrong tail.

Now Welch’s version. The standard test pools the two groups’ spreads into one SD, which assumes the two populations share a variance. Welch’s version lets each group keep its own variance, and pays for that with a fractional df. Run it and compare the two side by side:

Here the two versions barely differ: a standard error of $3.41 against $3.43, df of 903 against 861.2, \(t_\text{obs}\) of 8.62 against 8.57, the same cutoff to three decimals, and both p-values far below .001. That is because the groups are large and their SDs similar ($53.09 and $49.38). When group sizes and spreads differ, the two can part company, and Welch’s is the safer choice — which is why infer’s t_test() uses it by default; the Module’s Test 1 shows the two side by side. The standard version is the one the counting rule explains, which is why Video 2 drew it. In Activity 2.3 you’ll let t_test() run both: with var.equal = TRUE it gives the standard test, and without it Welch’s, so each output should match one line of your Activity 2.2 output, up to rounding in the last digit or two.

Activity 2.3 — Read the output 🔍

Nothing to write here. The chunk calls t_test() twice, once for each version you built by hand in Activity 2.2. Run it, then answer the questions below it before you click Show the answers.

Three questions, from the two printed rows:

  1. Which column is \(t_\text{obs}\)? Match each output to a row of your Activity 2.2 table: which is the standard test, and which is Welch’s? How do both compare to the \(t^*\) of 1.963?
  2. How does each p_value compare to \(\alpha = .05\)? Do the two versions reach the same decision?
  3. Do lower_ci and upper_ci contain zero, in either output?

The first call, with var.equal = TRUE, is the standard test from Video 2: statistic ≈ 8.62 on t_df = 903, matching the standard line of your Activity 2.2 output. The second is Welch’s, which t_test() runs unless you ask otherwise: ≈ 8.57 on t_df = 861.2, matching your Welch line. Each matches its line, though not to every digit: the p-values differ in their third digit, and Welch’s t_df is 861.22 here against 861.18 by hand. That is rounding, not a different test. In Activity 2.2 you typed the numbers off the table, rounded to the cent; t_test() works from the unrounded data.

Either way, the decision is the same. Both statistics are far past \(t^* \approx\) 1.963 — by a factor of more than four. Both p-values are far below .05. And both intervals sit entirely above zero: $22.68 to $36.06 for the standard test, the interval Video 2 pointed to, and $22.64 to $36.09 for Welch’s. The paper itself reports the Welch version, t(861) = 8.57, p < .001: your second line, which is why its df is 861 rather than 903.

Within each output, those are three readings of one decision, and they agree, by construction. The p-value, the cutoff comparison, and the interval are one statement written three ways (the Module’s on-ramp shows the algebra). So if these three ever disagree in a t-test’s output, something is mismatched: a one-sided test against a two-sided cutoff, mismatched df, or a confidence level that isn’t \(1 - \alpha\). It is a free correctness check.

Quick Check

Answer each question — you’ll see green (correct) or pink (incorrect) feedback as you type.

1. In this study, what is the estimand, the quantity the test is ultimately about?

2. The test is two-sided with α = .05. Where does that 5% go?

3. There are 905 participants. Why does the standard two-sample t use 903 degrees of freedom?

4. The p-value was about 3 × 10⁻¹⁷. What is that the probability of?

5. A two-sided t-test and its matching 95% t-interval are run on the same comparison. tobs exceeds t*, but the interval contains zero. What is the most likely explanation?


Video 3 — Cohen’s d and the magnitude question

What to listen for:

  • A p-value answers “how incompatible is this with the null model?” — and because incompatibility depends on precision, it shifts with sample size. Cohen’s d answers a different question: how big is the gap? The two are not competing answers; a significant result without an effect size is incomplete.
  • The construction is one small change from the test statistic: divide the same gap by the standard deviation (how much individual people differ) instead of the standard error (how much a sample mean wobbles). The same $29.37 gap gives \(t_\text{obs}\) = 8.62 over the $3.41 standard error, and d ≈ 0.57 over the $51.13 pooled SD. The SE shrinks as \(n\) grows; the SD does not.
  • What d looks like. At Cohen’s small, medium, and large benchmarks (d = 0.2, 0.5, 0.8), the two distributions still share about 92%, 80%, and 69% of their area. Our d ≈ 0.57 sits just past medium: a real gap, but far from clean separation.
  • The common-language effect size (CLES) turns d into a probability: pick one person from each group at random; how often does the person from the higher-scoring group score higher? No difference gives 50%, and it climbs as the means pull apart: about 56%, 64%, and 71% at d = 0.2, 0.5, and 0.8, the values in the subtitles of the video’s overlap panels. It assumes Normal curves with equal spread, and it is not the share of people who benefited.
  • Small, medium, and large are conventions, not measurements; prior work on your own question is the better yardstick. And lead with the raw difference: report the gap in Ice Dollars first, then add d, which is what lets you compare effects across measures and studies.

Activity 3.1 — Compute d ✍️

Video 2 found strong evidence against equal population means. Now: how big is the gap? You’ll compute Cohen’s d by hand first, then let the cohens_d() function from the effectsize package do it for you, and check that the two agree.

The only change from the test statistic is the denominator. Instead of the standard error — how much a sample mean wobbles — divide the gap by the pooled standard deviation itself: the two group SDs combined, each weighted by its degrees of freedom (the Going further box in Activity 2.1 shows the calculation). It’s the same s_pooled you used in Activity 2.1, printed below beside the standard error so you can see the swap.

Your task: one blank — the divisor that turns the gap into d.

Cohen’s d is the gap over the pooled SD, so s_pooled goes in the blank.

Your hand calculation and cohens_d() should agree: d ≈ 0.57, 95% CI [0.44, 0.71]. Note the pooled SD is about $51.13 — roughly fifteen times the $3.41 standard error from Activity 2.1. Same gap on top; a very different denominator underneath. That single swap is the whole difference between a test statistic and an effect size. When you write it up, lead with the raw $29.37, then add d: the raw gap says how big the difference is in units readers understand, and d puts it on a scale you can compare across measures and studies.

Sign watch: cohens_d() subtracts in factor-level order. graph_type lists CI before PI, so d is CI − PI and comes out positive; reverse the levels and only the sign flips.

The pooled SD assumes the two populations share a variance, the same assumption as the standard t you built in Activities 2.1 and 2.2. Welch’s test drops that assumption, and its matching effect size drops it too: pooled_sd = FALSE standardizes by \(\sqrt{(s_{CI}^2 + s_{PI}^2)/2}\).

Here the two agree to two decimals because the two SDs are so similar. The Module’s Test 1 sets out the rule (pooled d beside the standard test, pooled_sd = FALSE beside Welch’s); either way, name the standardizer you used.

The common-language effect size (CLES) for our data. Video 3 turned d into a question anyone can follow: pick one CI viewer and one PI viewer at random. How often does the CI viewer name the higher price? The effectsize package’s p_superiority() (“probability of superiority” is another name for the CLES) answers it straight from the data:

parametric = FALSE tells it to compare every CI viewer with every PI viewer and count how often the CI viewer names the higher price, with ties counted as half. The answer is about 68% (95% CI 65% to 71%). Without that argument it uses the conversion Video 3 described, which assumes Normal curves with equal spread and gives 66%; because willingness to pay is skewed and heaped on round numbers, counting is the safer choice here; report that one, and say which you used. (To build intuition for d itself, Kristoffer Magnusson’s interactive Interpreting Cohen’s d lets you slide d and watch the overlap and CLES change; set it near 0.57.) Either way, the CI display raises the chance of paying more from the 50% of no difference to about two in three, yet in about 28% of pairs the PI viewer still names the higher price. An average difference is entirely compatible with enormous individual variation, and that is invisible in the p-value.

Everything above used the matching-caption participants. The other half of Experiment 1 saw a caption that spelled out both intervals in words, whichever chart was on screen. That subset is pre-built as et; rerun d on it:

Here d ≈ 0.36, 95% CI [0.22, 0.49]: descriptively smaller than 0.57, and still comfortably clear of zero. Those are the two values the authors report, 0.57 and 0.36, so you have just reproduced them. Be careful how you say this. Comparing two separately estimated effect sizes does not test whether the caption changes the visualization’s effect; that is an interaction, and it needs a model that handles both factors at once, which you’ll meet in PSY 653. What you can say is narrower: the estimated CI−PI gap is smaller when the caption spells out both numbers, but it remains positive.

Activity 3.2 — Same d, different n 🔮

Predict first. Below, the effect is pinned at d = 0.57 and only the sample size changes. Before you run anything, commit to an answer:

As n per group goes from 30 to 100 to 450, what happens to (a) the t-statistic, (b) the p-value, and (c) Cohen’s d?

Write it down — actually write it down — then run the chunk.

At 30 per group \(t_\text{obs}\) is about 2.2 and p ≈ .03 — significant, but only just. At 450 per group \(t_\text{obs}\) is past 8 and p is vanishingly small. And d never moves, because we pinned it.

The formula shows why: \(t_\text{obs} = d \times \sqrt{n/2}\). The test statistic is the effect size multiplied by a function of sample size. They are not competing measures of the same thing — one is the effect, the other is the effect scaled by how well you measured it. (That formula is the equal-n, pooled-SD case; with unequal n the same principle holds.)

Two consequences worth carrying:

  • A tiny p-value does not mean a large effect. It can mean a modest effect measured very precisely — which is exactly our 905-participant study.
  • This is not a power simulation — we held the observed standardized gap fixed simply to isolate what changing n does to \(t_\text{obs}\) and p. Still, it shows why a non-significant result does not mean no effect: at n = 30 per group, an observed d of 0.57 only just clears .05, and a slightly smaller one would not have. That’s a statement about the study’s precision, not about the world.

Quick Check

Answer each question — you’ll see green (correct) or pink (incorrect) feedback as you type.

1. In the conventional two-sample Cohen’s d you computed in Activity 3.1, the gap between the means is divided by what?

2. Cohen called d = 0.8 “large.” What does a d of 0.8 mean?

3. At d = 0.8, the common-language effect size is about 71%. What does that 71% mean?

4. A colleague reports p < .001 and no effect size. What have they not told you?

5. Which write-up follows Video 3’s advice on reporting an effect?

6. In Activity 3.2, why does d stay fixed while the p-value gets smaller as n increases?


Video 4 — Where F comes from

A note on the data: this video and Activity 4.1 work on the Module’s data: Experiment 2, the four displays at the large effect size (N = 910). The numbers match the Module’s Test 2, and Video 5 returns to Experiment 1.

What to listen for:

  • Why not six t-tests? Four groups make six pairs, and every unadjusted test carries its own false-alarm risk. If six tests were independent, the chance of at least one false alarm would be 1 − .95⁶ ≈ 26%; ours share groups, so 26% is an illustration, not the exact rate, but the point holds. ANOVA asks one omnibus question first.
  • The hypotheses. \(H_0\): all four population means are equal. \(H_a\): at least one differs. The alternative does not claim that all four differ, and it does not say which one, by how much, or in which direction.
  • Total variation splits exactly into two pieces: variation of individuals around their own group’s mean (within), and variation of the group means around the grand mean (between). \(SS_\text{total} = SS_\text{between} + SS_\text{within}\) (2,861,857 = 60,946 + 2,800,910) is an identity, not an approximation.
  • You cannot compare the two sums directly; they are built from different numbers of free quantities. Divide each by its df (3 and 906) to get a mean square, and then \(F_\text{obs} = MS_\text{between} / MS_\text{within}\) = 20,315 / 3,092 = 6.57. When \(H_0\) is true both mean squares estimate the same population variance, so the ratio averages about 1 (most values land below 1, a few far above).
  • The F reference distribution is right-skewed, and the test uses only its right tail. F can’t be negative, and more separation among the means only makes it larger. With df = 3 and 906, \(F^*\) = 2.61; our 6.57 lands far past it, p ≈ .0002. The curve assumes independent observations, a common within-group variance, and Normally distributed errors; Welch’s ANOVA is the alternative when the variances differ a lot.
  • A significant F is a beginning. It says the four population means are not all equal, not which ones differ. Adjusted follow-ups answer “which” (Tukey’s method for every pair, planned contrasts for comparisons named in advance), and η² = \(SS_\text{between} / SS_\text{total}\) ≈ .021 says how large: about 2% of the variation goes with which display people saw.

Activity 4.1 — From sums of squares to F ✍️

This activity uses Video 4’s data: all_groups, the Module’s Experiment 2 at the large effect size, with the four displays in graph_type (CI, CI rescaled, PI, HOPS) and willingness to pay in wtp. The three sums of squares are computed for you, each directly from the data as the Module and the video build them, with the grand mean and the group means alongside. Run this as it stands:

Look at the relative sizes before you go on. Almost all the variation in what people were willing to pay is within conditions — person-to-person differences the experiment doesn’t explain. The between-groups piece is the small remainder tied to the four condition averages sitting apart — where a condition effect would show up, together with the little accidental separation any four random groups have. You cannot compare those two sums directly, though: one is built from 4 group means, the other from 910 observations. Divide each by its degrees of freedom — quantities minus constraints, the counting rule from Video 1 — and the two become comparable.

Your task: fill in the three blanks in the code below.

  1. df_between: the degrees of freedom for the between-groups piece, written with k, the number of groups.
  2. df_within: the degrees of freedom for the within-groups piece, written with N, the number of people, and k.
  3. f_obs: the mean square that goes on top of the F ratio.

The code then divides each sum of squares by its df, prints your F, and prints R’s own ANOVA table beneath it so you can check your work. That last line is the Module’s: lm() fits a model of willingness to pay by display, anova() builds the ANOVA table from it, and tidy(), from the broom package, turns the table into a tibble. From Module 10 on, that lm()-plus-broom workflow becomes your standard toolkit, and the t-test and ANOVA turn out to be special cases of the linear model it fits.

With k groups constrained by one grand mean, df_between is k - 1. With N observations constrained by k group means, df_within is N - k. And F puts the between mean square on top.

Your hand-built \(F_\text{obs}\) and R’s table agree: F(3, 906) = 6.57, p < .001.

Notice the reversal that dividing by df produced. \(SS_\text{within}\) (2,800,910) is about 46 times \(SS_\text{between}\) (60,946) — on the raw sums, the within-group variation dwarfs everything. But \(MS_\text{between}\) (20,315) is about 6.6 times \(MS_\text{within}\) (3,092). The sums of squares and the mean squares point in opposite directions, which is precisely why you cannot skip the df step.

So the four displays differ — but how much of the variation does display format account for? That is η²:

η² ≈ .021, 95% CI [.005, .041] — display format accounts for about 2.1% of the sample variation in willingness to pay. Set that next to p < .001 and the contrast is the whole lesson of Video 3 again: an effect can be nearly impossible to attribute to sampling variation and still be a small slice of what’s going on.

The F-test is omnibus: it tells us the four population means are not all equal, but not which means differ. Follow-up comparisons answer that next question, as Video 4 described: Tukey’s method when you want every pair, planned contrasts when you named specific comparisons in advance. The Going further box below runs Tukey’s. (Experiment 2 also varied the effect size. Like the Module, we use only the large-effect half, so display format is the single factor and a one-way ANOVA is the right model.)

Tukey-adjusted comparisons test every pair of displays while accounting for testing several pairs at once:

Read the pattern rather than the individual rows: only two pairs are statistically detectable after the adjustment, and both involve the CI display. People who saw CI error bars were willing to pay more than those who saw prediction intervals, and more than those who saw HOPS. The other four pairs, including CI against its rescaled version, are not distinguishable at these sample sizes. The Module’s Test 2 reads the same table.

One more point you may meet: Video 4 described the most common order, where you run the omnibus test first and, if it is significant, follow up with adjusted comparisons. That order is a habit, though, not a requirement. Tukey-adjusted comparisons control their own familywise error rate, and planned contrasts specified in advance can be tested regardless of how F came out. The Module’s Test 2 sets this out.

Video 4 says that when \(H_0\) is true, F averages about 1, with most values below 1 and a long right tail. You can check that by simulation. The chunk below draws four groups of 228 (about the size of each display group) from a single population, so there is no real difference to find, computes F, and repeats that 300 times:

The average lands close to 1 (0.98), about 64% of the values fall below 1, and the histogram is right-skewed: F can’t go below zero but has no upper limit. About 5% land past the dashed \(F^*\) line. Those are false alarms from a world with no effect, which is α doing exactly what it advertises. Against that backdrop, our \(F_\text{obs}\) of 6.57 is in a different regime entirely.

Quick Check

Answer each question — you’ll see green (correct) or pink (incorrect) feedback as you type.

1. Why does Video 4 start with one ANOVA instead of six pairwise t-tests?

2. Which statement matches the alternative hypothesis for this ANOVA?

3. Why must sums of squares be converted to mean squares before they are compared?

4. When H₀ is true, F averages about 1. Why 1 rather than 0?

5. Why does the F-test look only at the right tail of its reference distribution?

6. Our ANOVA gave p < .001 with η² ≈ .021. What is the right read?

7. True or false: because F(3, 906) = 6.57 with p < .001, we now know that every pair of displays differs — for example, CI rescaled versus HOPS.


Video 5 — When the outcome is categorical

What to listen for:

  • Same two groups, new kind of outcome. The 905 people from Videos 2 and 3 come back, but the question is now yes/no: was each person willing to pay more than the upgrade is worth — the $17.50 risk-neutral price? Two displays by two answers makes a 2 × 2 table, and the test asks whether the share who overpaid is the same under both displays.
  • Two symbols run every chi-square test: O, the count you observed, and E, the count expected if the null were true. In a test of independence, the margins do double duty: they generate the expected counts (row total × column total ÷ overall total), and they fix the degrees of freedom — in a 2 × 2, fill in one cell and the other three are forced, so df = 1.
  • Why the departures are squared. In a 2 × 2 every cell misses its expectation by the same amount — two cells above, two below — so the raw departures cancel to exactly zero. Squaring makes every cell count, and dividing by E scales each miss to the size of the cell it came from. It is also why the chi-square statistic itself has no direction.
  • Compared to what? The reference curve is a chi-square with 1 df, piled up against zero, with only an upper tail, because squared departures can only push the statistic up. The α = .05 cutoff is 3.84; our 31.81 lands far past it, p < .001.
  • Check the expected counts, not the observed ones. The chi-square curve is a good approximation only when the expected counts are large enough: none below 1, and no more than about 20% below 5. Here the smallest is 81.5, so every expected count is comfortably above 5 and the approximation is excellent.
  • Evidence is not magnitude. The raw effect is the two row percentages, 89% of CI viewers against 74% of PI viewers. Cramér’s V (for a 2 × 2, \(\sqrt{\chi^2 / N}\), also called the phi coefficient) says how strong the association is: about .19, a small one. With 905 people, a modest gap still gives an emphatic χ².

Before we cut a continuous variable in two

Willingness to pay is measured in dollars. Turning it into a yes/no throws information away, and we are doing it here to practice the categorical test, not because it improves on the analyses you just ran.

A defensible reason to dichotomize is a substantively meaningful cutpoint — and Experiment 1 supplies one. The upgrade’s risk-neutral price is $17.50: that is the most it is worth in expected winnings. It comes from the boulder itself: in Experiment 1 the special boulder wins 57% of the time, so renting it is worth \(250 \times .57 - 250 \times .50 =\) $17.50 in expected winnings, the same calculation the Module uses. Naming a higher price means offering more than a risk-neutral player should — which is a precise, model-based statement, not a verdict that the person was irrational. Throughout this section, read “overpaid” as shorthand for above the risk-neutral benchmark.

Cutting at a convenient round number instead ($50, say, because it looks like a reasonable slice of the $250 prize) is how dichotomization earned its bad reputation. The number has to come from the problem, not from the keyboard.

Activity 5.1 — Build the cross-tab ✍️

You built cross-tabs in M05; the only new thing here is where the cut goes. We stay with mt, the same 905 people as Videos 2 and 3, so the rows are the two displays you already know.

Your task: fill in the three blanks in the code below.

  • Blanks 1 and 2: the cut that splits the outcome, the risk-neutral price. It goes in both lines of the case_when(): above it is overpaying, at or below it is not.
  • Blank 3: the percentage direction that makes the null hypothesis easy to see: each display should sum to 100%.

The cut is at the risk-neutral price, 17.50, and it goes in both conditions: strictly above it is overpaying, at or below it is not. For the percentages, you want each display to sum to 100%, so the direction is "row".

Overall, 81% of participants were willing to pay above the risk-neutral benchmark. But read across the two rows: 89% of CI viewers overpaid, against 74% of PI viewers. That is the null hypothesis made visible. Independence is the claim that, in the population, the two proportions are equal — that knowing which display someone saw tells you nothing about whether they overpaid. Our two sample percentages are clearly not equal; the test asks whether they are further apart than sampling variation would ordinarily produce.

Activity 5.2 — Expected counts, then the test ✍️

This activity has three parts: build the expected counts, run the test, and measure the effect. Each uses something new, so here is how each piece works before you run it.

1. The expected counts. Each cell’s expected count is its row total × column total ÷ overall total, so the code first writes those totals next to every cell, then multiplies. It takes four steps, each adding one column you can see:

  • Step 1: count() makes one row per cell (CI and did not overpay, CI and overpaid, and so on), with the number of people observed in it. You used count() in M04; the new piece is name = "observed", which labels the count column observed instead of count()’s default, n.
  • Steps 2 and 3 are M04’s grouped mutate. group_by(graph_type) + mutate(row_total = sum(observed)) adds up the counts within each display and writes that total on both of the display’s rows, keeping every row (summarize() would collapse them). ungroup() clears the grouping so the next step starts fresh. Step 3 repeats the move by overpaid to get the column totals.
  • Step 4: one mutate() does the arithmetic.

These expected counts are what the adequacy guideline applies to, not the counts you observed.

2. The test. chisq_test() comes from infer, the same package as the t_test() you ran in Activity 2.3, and it reads the same way: pipe in the data, then give formula = outcome ~ group. correct = FALSE switches off an adjustment R applies to 2 × 2 tables only (Yates’s correction), so the statistic matches the one built from your expected counts. It returns one row: statistic is \(\chi^2_\text{obs}\), chisq_df is its degrees of freedom, and p_value is the area beyond it on the chi-square curve.

3. The effect size. cramers_v(), from effectsize, wants the 2 × 2 table of counts itself rather than a formula. So select() picks the two variables and table() cross-tabulates them. adjust = FALSE asks for the plain V, \(\sqrt{\chi^2 / N}\) (the default shrinks it slightly to correct for small samples), and alternative = "two.sided" asks for an ordinary two-sided 95% CI (the default reports a one-sided interval capped at 1).

Your task: fill in the three blanks in the code below.

  • Blank 1: the overall total that finishes the expected-count formula (row total × column total ÷ overall total).
  • Blanks 2 and 3: the two variables in the test’s formula: the yes/no outcome on the left of the ~, the display on the right.

The overall total is the sum of all four observed counts, sum(observed): 905 people. (After Step 3’s ungroup(), sum() adds up the whole column, not one group.) The formula puts the outcome first and the grouping variable second, exactly as t_test() did: overpaid ~ graph_type.

χ²(1, N = 905) = 31.81, p < .001, Cramér’s V ≈ .19, 95% CI [.12, .25].

Three things to read:

The expected counts tell the story before the test does. Compare each observed with its expected: every cell misses by the same 33.5 people — two cells above expectation, two below. That is the margins at work (whatever one cell gains, its row-mate and column-mate must lose), and it is why a 2 × 2 table has exactly one degree of freedom. The smallest expected count is 81.5, well clear of the guideline, so the chi-square curve is a reliable reference here.

The statistic is the one the video built. Square each departure, divide by its expected count, add the four: 31.81. That is what correct = FALSE buys you — leave it out and R quietly applies a 2 × 2-only adjustment and reports 30.87 instead. Either is defensible; the Module’s Test 3 explains the difference, and the habit to build is saying which one you used.

The effect is small. The raw effect is the pair of percentages — 89% against 74% — and Cramér’s V of .19 (for a 2 × 2, the same number as the phi coefficient, \(\sqrt{\chi^2 / N}\)) puts it at the small end of the conventional 2 × 2 benchmarks. Set that beside p < .001 and you have the evidence-versus-magnitude habit one more time: with 905 people, a modest difference in proportions produces an emphatic χ². And notice that, with one degree of freedom, no follow-up is needed to say where the association lies — the two row percentages already say it. It is the same CI-over-PI gap you measured with t and d in Activities 2.1 through 3.1, re-expressed as a yes/no, so it echoes that result rather than independently confirming it.

Quick Check

Answer each question — you’ll see green (correct) or pink (incorrect) feedback as you type.

1. In a chi-square test of independence, where do the expected counts come from?

2. Why does Pearson’s chi-square square the O − E deviations rather than simply adding them up?

3. In a 2 × 2 table with its row and column totals fixed, why is there exactly one degree of freedom?

4. Before trusting a chi-square test of independence, what should you check?

5. Our test gave χ²(1, N = 905) = 31.81, p < .001, with Cramér’s V ≈ .19. What is the right read?


Wrap-up

Four things to carry into lecture

  1. One logic, different designs. Every test asks whether the observed departure is large relative to what the null model would ordinarily produce. You built the two pieces of \(t_\text{obs}\) by hand, then \(F_\text{obs}\) from mean squares, and the shape was the same both times — and the three readings of a test (statistic against critical value, p against α, interval against the null value) agree by construction.
  2. Evidence is not magnitude. A p-value addresses compatibility with \(H_0\); a raw effect and an effect size say how large the difference is. Pinning d at 0.57 while dialing n took the p-value from about .03 to far below .001. Report both, and lead with the raw units.
  3. ANOVA is signal over noise. F compares between-group variation with within-group variation. Near 1 looks like the null; a large \(F_\text{obs}\) says the means are not all equal — not which ones; follow-up comparisons say which.
  4. Chi-square compares observed with expected. Under independence, the margins tell us what counts to expect. χ² gets large when the observed table departs strongly from those expectations — and the expected counts are also what tells you whether to trust the test. In a 2 × 2, the two row percentages are the raw effect, and Cramér’s V is its size.

Before running any test, ask: What is my outcome? What groups or measurements do I have? Are they independent or paired? What parameter am I trying to learn? — the decision map from the Module, and the four questions every analysis in this course starts with.

And there is light at the end of the tunnel. The t-tests and the ANOVA you practiced here are special cases of a single model, the linear model, which Part 3 builds with lm(). Learn it, and this list of tests becomes one formula you adjust.