Reading Statistics #1: Stop Interpreting Papers Using 'Significant Differences' — P-values and 95% Confidence Intervals
"Reading Statistics" is a series on medical statistics and data science for clinicians and medical researchers. It is not about the technique of "calculating" statistics, but about honing the ability to "interpret" them. In this first installment, we will interpret the two most familiar and most carelessly read numbers: P-values and 95% confidence intervals.
Estimated reading time: approx. 6 minutes / Prerequisites: None
Two papers on the desk
There are two papers on your desk.
One has "statistically significant (P < 0.001)" written in large letters.
The other is summarized as "no significant difference (P = 0.10), negative trial."
Which treatment should you believe?
Many people would choose the former without hesitation.
However, that judgment is likely wrong.
This is because the magnitude of a P-value is not a ranking of the quality of a treatment.
Why can we say that? First, let's start with what a P-value actually means.
Key Points
A P-value is neither the "probability that there is no effect" nor the "probability that the result is due to chance." It is merely the probability of obtaining data as extreme as, or more extreme than, what was observed, assuming the null hypothesis is true.
The binary evaluation of "significant difference / no significant difference" discards most of the information for judgment. P = 0.049 and P = 0.051 are almost identical evidence, and the boundary of 0.05 has no special clinical meaning.
What should be read as the effect is the confidence interval. How far the interval extends toward the null hypothesis value and the "clinically meaningful value"—that is where it is revealed what the study could and could not say.
1. What a P-value is, and what it is not
A P-value is "the probability of obtaining data as extreme as, or more extreme than, what was observed, assuming the null hypothesis is true." The null hypothesis often used here is that "the treatment has no effect."
That is the only definition. And from this, three "common misinterpretations" arise.
First, a P-value is not the "probability that there is no effect." P = 0.10 does not mean there is a 10% probability that there is no effect. A P-value is a number calculated under the assumption from the start that "there is no effect," and it does not measure the correctness of that assumption itself. Second, a P-value is not the "probability that the result is due to chance." This is just another way of stating the same error. Third, the number obtained by subtracting the P-value from 1 is not the "probability that the treatment is effective." The interpretation that P = 0.04 means there is a 96% probability that the treatment effect is real is invalid and incorrect.
A P-value answers only one thing: "If there were no effect, would data like this be rare or common?" That is all.
Therefore, a large P-value (not significant) does not mean 'it has been shown that there is no effect,' but simply that 'even under the assumption that there is no effect, this data was not particularly unusual.'
This is the reason why the 'negative trial' mentioned at the beginning does not constitute 'proof of no effect.'
2. What 'Significant/Not Significant' Overlooks
The next problem lies in the line of 0.05 itself.
If P=0.049, it is 'significant'; if P=0.051, it is 'not significant.' We treat these two as if they were black and white. However, the difference between the two as evidence of a treatment effect is almost non-existent. The value of 0.05 is a convention in the history of statistics, not a boundary line that separates truths existing in the natural world.
The harmful effects of binary evaluation have been warned about by experts themselves many times. In 2016, the American Statistical Association, and in 2019, hundreds of researchers in the journal Nature, publicly called for an end to 'relying on binary evaluation called significant difference.' Those at the center of statistics are the very ones calling a clear 'halt' to the use of statistics' most famous tool.
What is even more troublesome is the opposite. 'Significant difference' also guarantees nothing on its own.
If you increase the sample size, no matter how small the difference, it will eventually become 'significant.' If you collect tens of thousands of cases, even a clinically meaningless difference will get an asterisk.
In other words, 'significant difference' does not mean 'large effect,' and 'no significant difference' does not automatically mean 'no effect.' We are entrusting our judgments to a result that is only one line long and says almost nothing.
3. Therefore, Read the Confidence Interval
So, where should we look? It is the 95% confidence interval.
The confidence interval shows the 'possible range' of effects compatible with the data. While the P-value only returns a single judgment of 'significant or not,' the confidence interval tells us two things at once.
The magnitude of the effect (where the interval is located) and its uncertainty (how wide the interval is).
If the interval is narrow, the study says a lot. If the interval is wide, even if a conclusion is reached, there is actually not much information. Most 'not significant' studies did not deny the effect; they just had intervals that were too wide to say anything at all.
And the real way to read a confidence interval is not to see whether it straddles the null hypothesis value (generally a value representing no difference; 1 for ratios, 0 for differences). That is just a roundabout way of checking 'significant or not,' and is no different from looking at a P-value.
What should be read is the width of the confidence interval expressed by a single line.
4. Compare with Clinically Important Values, Not the Null Hypothesis Value
Compare the interval with the 'magnitude of a clinically meaningful effect.' This is the core.
Let's look at two examples. Both use fictional numbers.
The first example is a study where the difference between groups is
Hazard ratio 0.88, 95% confidence interval 0.75–1.03, P=0.10
This is the kind of paper that might be dismissed on your desk as 'not statistically significant, negative.' But if you read the interval, the data is consistent with anything from a '25% risk reduction' to a '3% risk increase.' This does not show that there is 'no effect.'
Even a clinically significant effect, such as a 25% risk reduction, has not yet been ruled out. The correct interpretation is not 'no effect,' but rather 'no conclusion reached.'
As a second example, in another trial:
Hazard ratio 0.98, 95% confidence interval 0.97–0.99, P<0.001
Statistically, it is undeniably significant. However, the magnitude of the effect is a mere 2% in relative risk. From one end of the interval to the other, it falls within a range that is clinically almost negligible. Concluding that it is clinically important just by looking at the word 'significant' is a misreading that is the opposite of the first example, but for the same reason.
Let's return to the two papers on your desk. The one with 'P<0.001' might be 'significant but trivial,' like the second example. The one with 'P=0.10' might be 'inconclusive, but with the possibility of a large effect,' like the first example. By looking only at the magnitude of the P-value, it is theoretically impossible to know which is correct until you read the interval.
Takeaways for today:
Stop reading asterisks (whether significant or not significant).
Read confidence intervals by comparing them against both the null hypothesis value and clinically meaningful values.
'Not significant' often means 'no conclusion reached yet,' not 'no effect.'
Next time, in #2, we will cover situations where 'significance' appears in more subtle forms. We will discuss the problem of multiplicity, where accidental significance almost inevitably arises as you perform multiple analyses, and the pitfalls of subgroup analysis.
