An annotated example report
Project 2 · a competent reproduction, marked up with what would raise it
How to read this page
Below is a short example reproduction report on an invented paper, written the way a competent team often writes a first draft. It is not a model answer — it is a competent one, which is where most first drafts land.
After each section you will find a boxed note — What would raise it — naming what separates competent from strong. The notes are the point of this page. This project’s rubric rewards characterizing a discrepancy rather than smoothing it over — the notes are what that difference looks like on the page.
The paper, the numbers, and the discrepancy are fabricated for illustration. Do not cite them.
The Project 1 annotated example does the same thing for the data brief.
The example report
1 · The published study and target
The claim
Marchetti and Okonkwo (2021) asked whether undergraduates who completed a brief sleep-hygiene workshop reported lower daytime fatigue than a waitlist group. They report that the workshop group scored lower on a fatigue scale at follow-up, t(178) = 2.61, p = .010, d = 0.39.
What would raise it
The claim is reported accurately and the statistics are copied correctly, which is the requirement. What is missing is the scope of the claim: the paper’s own conclusion was about a workshop delivered to volunteers at one university, and saying so here sets up the discussion in the final section. State the claim at the size the authors actually made it.
The section also never says why the finding matters. A sentence placing it in its research area (how common daytime fatigue is among undergraduates, say, and whether brief workshops are a realistic way to reach them) tells the reader why a reproduction is worth their time.
The design, the estimand, and the hypotheses
Two independent groups, one continuous outcome, measured once. That is an independent-samples design, so the estimand is the difference in population mean fatigue between workshop and waitlist, and H0 is that the difference is zero.
What would raise it
Correct, and stated in the team’s own words rather than copied — which is what the rubric asks. M09’s template names the estimand in words and then in symbols (\(\mu_\text{workshop} - \mu_\text{waitlist}\)), and the draft stops after the words. A stronger version would say why the design is independent rather than paired here, since the paper measured everyone at follow-up only. Naming what the design is not shows you recognized it rather than inherited it.
2 · Data and analytic-sample reconstruction
The analytic sample
The posted dataset has 214 rows. The paper reports 180 in the analysis. Applying the exclusions described in the Method — attrition before follow-up, and one group assignment recorded as unknown — leaves us with 186.
What would raise it
This is the most valuable paragraph in the report and it stops one sentence early. A six-person gap is a finding, not an inconvenience: which exclusion is ambiguous, and what would settle it? A strong version says “the paper does not state whether partial-completion cases were dropped; if we also exclude the four participants missing two or more scale items, we reach 182, and at 180 we would additionally need a criterion the paper does not describe.” That is a reader who could act on your work.
The descriptives, and the choice of test
Group means and SDs are close to the published table. A Q-Q plot of the outcome shows mild right skew in both groups, and group variances are similar, so we used Welch’s t-test.
What would raise it
The assumption check is present and the choice of Welch is defensible. Two things would raise it. First, the mild skew is reported but never interpreted: at about 90 people per group, M08 and M09 say it barely affects the t-test, and saying so is the judgment the rubric rewards. Second, “close to the published table” is doing a lot of unexamined work — how close? Reporting the two sets of means side by side, before the test, is what lets a reader see whether a later discrepancy came from the sample or the analysis. The rubric asks for the descriptive reconstruction before the test for exactly this reason.
The version of the test should also follow the paper, not the variances alone. A stronger report runs both the standard and Welch’s two-sample t, checks which one’s df matches what the authors printed, and says so.
3 · Reproduction analysis · 4 · Original versus reproduced result
[Figure 1, titled “Fatigue by group”: each participant’s follow-up fatigue score, with the group means and 95% CIs.]
| Quantity | Original | Ours | Match |
|---|---|---|---|
| Analytic N | 180 | 186 | close |
| Test statistic | t(178) = 2.61 | t(181.4) = 2.38 | close |
| p-value | .010 | .018 | same direction |
| Effect size | d = 0.39 | d = 0.35 | close |
Our reproduction found the same direction and a slightly smaller effect.
What would raise it
The table is the right deliverable, built the way the rubric asks, but its Match column is too vague to do its job. “Close” for an analytic N of 180 against 186 hides the discrepancy that matters most here, a different sample. And “same direction” for the p-value is the same-side-of-.05 comparison the project page warns against. Describe each match specifically (exact, rounding, or differs), and name the likely reason for anything that differs. The single sentence under the table is where the report says too little. Note that the df differ in kind, not just size: 178 is exactly the authors’ N − 2, which points to the standard (pooled) t, while 181.4 can only be Welch’s. That means the team probably did not run the same test the authors ran, which is a legitimate choice but has to be said, because it is one candidate explanation for the p-value moving.
Three required pieces are thin or missing here:
- The raw effect. Nothing reports the difference in mean fatigue or its 95% CI, which M09 puts first, before any standardized number. The table needs that row, and a row for the group means.
- The standardizer. The team ran Welch’s, so its d should use the unpooled SD; the report doesn’t say which one d = 0.35 used.
- The classification. The page asks every team to name the outcome. Here the N differs beyond rounding and the gap is described, so this is a partial reproduction, and the report should say so in those words.
The figure is the right kind of figure, but its title names the variables rather than the finding. “Workshop participants reported less fatigue than the waitlist” does the reader’s work for them.
5 · Discussion and scope
The finding reproduced in direction and approximate magnitude. The p-value moved from .010 to .018, likely because our analytic sample is slightly larger. This does not undermine the original result. Limitations include our uncertainty about the exclusions and the fact that we used Welch’s test.
What would raise it
The honesty is right and the team resists overclaiming, which many drafts fail to do. But “likely because our analytic sample is slightly larger” is a guess presented as a conclusion — and it is checkable. Running the test again on the 180 closest to the paper’s description, and reporting whether p returns to .010, converts a guess into evidence. That single extra run is the difference between a team that noticed a discrepancy and a team that characterized one.
Two more things. “Reproduced in direction and approximate magnitude” falls back on direction, which is not one of the page’s criteria; the classification above says it better. And “this does not undermine the original result” claims more than the team has shown, since the sample gap is unresolved.
The section is also titled Discussion and scope, but it never says what the design permits. Were participants randomly assigned to the workshop or the waitlist? That one fact decides whether “the workshop reduced fatigue” is a causal claim or only a difference between groups.
The last sentence also lists limitations rather than weighing them. Which of the two matters more for whether the original finding stands?
What this example is not
It does not show a strong report, deliberately. The distance from competent to strong here is not more analysis — it is a handful of specific moves:
- reporting the claim at the size the authors made it
- treating a sample-size gap as a finding to diagnose, not a nuisance to note
- showing the descriptive comparison before the test, so a later gap can be located
- reporting the raw effect and its interval first, and a figure whose title states the finding
- saying when you ran a different test from the authors’, and why that matters
- naming the outcome with the page’s categories (near-exact, partial, or not reproduced) and saying what the design permits
- checking the explanation you offer, when checking it costs one more line of code
The last one is the whole project in miniature. A reproduction that says “probably because…” has stopped one step short of its own contribution.