Calcylator
P Value

The p-value:
how surprising is the result if nothing is going on?

The most quoted number in research, and one of the most commonly misread. A plain walk through what it measures.

Calcylator Editorial Team

Updated · 5 min read

The recipe: statistic, distribution, tail

Suppose someone claims a coin is fair and you flip it 100 times, getting 60 heads. Is that suspicious? The p-value answers a narrow question: if the coin really were fair, how often would a result at least this far from 50 heads turn up by luck alone? If the answer is 4.6% of the time, the data are somewhat surprising under the fair-coin story.

Everything hangs on the phrase assuming the null. The p-value starts from a model with no effect and measures how well your data fit it. A small p-value says the data fit badly; it does not say how large, how important or how likely the alternative is.

P-value (two-sided) =p = P(|Z| ≥ |z observed|) under the null
Z:
the test statistic's null distribution (standard normal here)
z observed:
the value computed from your data
|·|:
absolute value, counting both tails
  1. State the null hypothesis, such as a fair coin or no difference between groups.
  2. Compute a test statistic from the data, such as z or t.
  3. Find how much probability lies at or beyond that statistic under the null.
  4. Double it for a two-sided test, which counts extremes in both directions.

Worked example: 60 heads in 100 flips

  • Null

    fair coin, P(heads) = 0.5

  • Expected heads

    100 × 0.5 = 50

  • Standard deviation

    √(100 × 0.5 × 0.5) = 5

  • z

    (60 − 50) ÷ 5 = 2.00

Two-sided p-value

p ≈ 0.046

The tail beyond z = 2 is 0.0228 on each side, so p = 2 × 0.0228 = 0.0455. This uses the normal approximation without a continuity correction.

Because 0.046 falls just under the conventional 0.05 line, many analysts would call it statistically significant. The coin is not proven unfair, though; it is only less compatible with fairness than with a slight bias.

A z-score to p-value reference

z (two-sided)p-valueEveryday reading
1.280.200Ordinary noise, happens 1 time in 5
1.6450.100Weak evidence
1.960.050The traditional line; 1 time in 20
2.580.010Strong evidence against the null
3.290.001Very strong; 1 time in 1,000

The famous 1.96 corresponds to 95% of a normal distribution sitting within it, which is why 0.05 and 95% confidence intervals go hand in hand. A one-sided test uses only one tail, so the same z of 1.96 gives p ≈ 0.025.

One-sided versus two-sided

A two-sided test asks whether the effect differs from the null in either direction. A one-sided test asks only about one direction and must be chosen before seeing the data. Switching to one-sided afterwards, because it halves the p-value, is a form of cheating.

A coin result of z = 2.40 is p = 0.016 two-sided and p = 0.008 one-sided. If you did not decide on a direction in advance, quote the two-sided value.

P-values and confidence intervals together

A confidence interval gives the range of effect sizes compatible with the data at a stated level, while the p-value gives a single number for compatibility with exactly zero effect. The two are linked: a two-sided p below 0.05 corresponds to a 95% interval that excludes the null value.

Take the coin example. The observed share of heads is 0.60 with a standard error of √(0.6 × 0.4 ÷ 100) = 0.049, so a 95% interval is 0.60 ± 1.96 × 0.049, or roughly 0.50 to 0.70. The lower end just touches 0.50, which is why p is close to 0.05. The interval is more informative because it also shows that biases of 0.55 or 0.68 are plausible.

Reporting only that p < 0.05 discards that information. Readers can tell whether an effect is large enough to care about only from the interval or the effect size.

Multiple comparisons and p-hacking

A 5% threshold means that if nothing is going on, one test in twenty will still cross it. Run twenty independent tests of true nulls and the chance that at least one reaches p < 0.05 is 1 − 0.95²⁰ = 64%. A researcher who tries many outcomes, subgroups or analysis choices and reports only the successes will almost always find something.

  • Pre-register the hypothesis and the analysis plan before collecting data.
  • Correct for the number of tests, for example with a Bonferroni adjustment that divides the threshold by the number of comparisons.
  • Report every test that was run, not just the significant ones.
  • Treat a surprising single result as a hypothesis to replicate, not a finding to build on.

Reporting a p-value honestly

Good practice is to give the exact value, such as p = 0.046, rather than p < 0.05, and to say which test and which tail were used. Reporting p = 0.000 is a rounding artefact; write p < 0.001 instead. State the sample size and the effect with its interval so that a reader can judge size and precision.

Treat 0.05 as a convention, not a cliff. A result at 0.049 and one at 0.051 carry almost the same evidence, yet one gets published as a finding and the other as a null. Think of the p-value as a continuous measure of how awkward the data are for the null model, and weigh it alongside study design, prior knowledge and replication.

What a p-value does not say

  • It is not the probability that the null hypothesis is true. Under a 0.03 result the null could still be quite likely, depending on prior plausibility.
  • It is not the probability that the result happened by chance in the sense of being a fluke.
  • It does not measure effect size. With a very large sample a trivial difference can give a tiny p-value.
  • A p-value above 0.05 does not prove no effect; it may only mean the study was too small to detect one.
  • Testing many hypotheses and reporting the one with p < 0.05 inflates false positives.

Common questions

What does a p-value of 0.05 mean?

It means that if the null hypothesis were true, results at least as extreme as yours would occur about 5% of the time, or 1 in 20. It is a measure of compatibility with the null, not proof of anything.

How do you calculate a p-value from a z-score?

Find the area in the tail of the standard normal curve beyond your z, then double it for a two-sided test. For z = 1.96 the tail is 0.025, so p = 0.050. Software or a z-table gives the tail area.

Is a p-value the probability the null hypothesis is true?

No. It assumes the null is true and asks how surprising the data are. The probability that the null is true is a different quantity that also depends on prior evidence and requires a Bayesian analysis.

What is a good p-value?

There is no single good value. A threshold of 0.05 is convention, and fields like particle physics demand far smaller values. Smaller p-values signal greater incompatibility with the null, but say nothing about effect size or importance.

Why can a tiny effect have a tiny p-value?

With a large sample the standard error shrinks, so even a trivially small difference becomes statistically distinguishable from zero. Always check effect size or confidence interval to judge practical importance.

Was this guide helpful?

Continue reading

View all blogs