THE PARSE / ISSUE 012

What a p-value does not mean.

2026-07-24 // BY Mira Chen // PHD, STATISTICS LEAD

In statistics coursework the software produces a correct number and the student writes a wrong sentence about it. The sentence is what gets marked, and there are only about four wrong sentences, which makes this unusually fixable.

001What it actually is.

A p-value is the probability of observing data at least as extreme as yours, assuming the null hypothesis is true. Every word in that definition is load-bearing, and the assumption at the end is the one that gets dropped in translation.

Notice what it is a probability about. It is a statement about the data, conditional on a hypothesis. It is not a statement about the hypothesis, conditional on the data, which is what nearly everyone wants it to be and what almost every incorrect interpretation quietly converts it into.

That single reversal generates most of the errors below. Once you can feel the difference between the probability of the data given the hypothesis and the probability of the hypothesis given the data, the wrong sentences start sounding wrong on their own.

002The four wrong sentences.

Each of these appears constantly in coursework, and each is marked wherever a grader is paying attention.

003What to write instead.

The correct sentence is less satisfying than the incorrect one and takes about the same number of words. Report the test, the statistic, the degrees of freedom, the p-value and the effect size in the format your course requires. Then state the conclusion in terms of the decision rather than in terms of truth: the null hypothesis was rejected at the stated level, and the observed difference was of a given magnitude.

Then, and this is the part that earns the marks, say what it means in the units of the actual problem. Not that the relationship was significant, but that a one-unit increase in the predictor corresponded to a change of a stated size in the outcome, and whether a change of that size is consequential in the setting being studied.

That last move is where the difference between bands usually sits. It is also the move that requires you to have thought about the subject rather than about the software, which is precisely why rubrics reward it.

004Assumptions, and the second half of the mark.

Most graduate rubrics carry a row for assumption checking, and it is one of the most commonly skipped rows in quantitative coursework. Normality, homogeneity of variance, independence, linearity: which ones apply depends on the test, and stating that you checked them is not the same as showing what the check found.

The honest version includes what to do when an assumption fails, which is where students freeze. An assumption violation is not a failure of the assignment; it is a finding, and the correct response is to say so, explain what it implies for interpretation, and either switch to a test that tolerates the violation or state the limitation explicitly. Graders reward that considerably more than a paper that ran the test anyway and said nothing.

Where the analysis belongs to a thesis rather than a course, this compounds, because a committee will ask about assumptions in a defense and a written record that ignored them is difficult to defend live. Analysis and the language describing it are produced together at the long-form desk for that reason.

005Why this is the highest-return concept in the sequence.

Statistics is cumulative in a way most subjects are not. The interpretation habits you form in an introductory course reappear in methods, in your thesis, and in the defense, and a misunderstanding that costs two marks in week four costs considerably more later.

It is also the most systematic error in the field, which makes it correctable in a single conversation rather than through general study. If your practice work keeps landing in the same place, the fault is far more likely to be a repeated misreading than a knowledge gap, and those need opposite remedies. Working through actual wrong answers with a statistician is what separates them, which is set out at the statistics desk.

Questions.

Should I report exact p-values or thresholds?

Exact values, to the precision your style guide specifies, unless the value is smaller than the guide's floor in which case use the conventional notation for that. Reporting only that a result cleared a threshold throws away information a reader needs, and most current style guidance asks for the exact figure.

Which effect size should I use?

It depends on the test: standardized mean differences for comparisons of two groups, variance-explained measures for analysis of variance, and the coefficients themselves for regression, interpreted in the units of the variables. Follow your course's convention where it states one, and always interpret the number rather than only reporting it.

My result was not significant. Does that ruin my paper?

No, and treating it as a failure is itself a mistake graders notice. A non-significant result is a result, and the correct write-up says what was not found, considers whether the study had enough power to detect an effect of a plausible size, and discusses what that implies. Papers that apologize for null findings score worse than papers that examine them.

How do I interpret a regression coefficient properly?

In the units of the variables, and in a sentence a non-statistician could follow. A one-unit change in the predictor is associated with a change of the estimated size in the outcome, holding the other variables constant. Then say whether a change of that magnitude means anything in the setting. That final clause is what most rubrics are actually looking for.

Mira Chen
WRITTEN BY
Mira Chen
PhD, statistics lead · one of eight leads on the bench.
THE PARSE // WEEKLY

One of these in your inbox each week.

SUBSCRIBE_FREE →
SEE_ALSO
Take my statistics classTake my psychology classCapstone and dissertationLiterature review structure
Send the course. We will tell you straight.GET_QUOTE →
ONLINE