What Is a P-Value? A Plain-English Explanation With Examples

A p-value measures how surprising your data would be if nothing interesting were going on. Learn how to calculate it, read it and avoid the classic misinterpretations.

A p-value is the probability of getting results at least as extreme as the ones you observed, assuming the null hypothesis is true. A small p-value, such as 0.02, means your data would be surprising if nothing were going on, which counts as evidence against the null hypothesis. A large p-value means the data are unremarkable under the null. It is not the probability that the null hypothesis is true.

That definition is precise and easy to misread, so this guide builds it up with examples before turning to interpretation.

The idea in one example

You flip a coin 100 times and get 60 heads. Is the coin biased?

  1. State the null hypothesis (H₀): the coin is fair, so the probability of heads is 0.5.
  2. State the alternative (H₁): the coin is not fair.
  3. Ask: if the coin really were fair, how often would 100 flips give a result at least as lopsided as 60 heads, in either direction?

Using the binomial distribution, the chance of 60 or more heads from a fair coin is about 0.0284, and the chance of 40 or fewer is the same. Together:

p = P(≥ 60 heads) + P(≤ 40 heads) ≈ 0.0284 + 0.0284 ≈ 0.057

A fair coin would produce a split this uneven about 5.7% of the time. That is somewhat unusual but not rare enough to clear the conventional 0.05 bar, so you would not reject the null hypothesis. The binomial probability calculator computes these tail probabilities exactly.

How to calculate a p-value

Every p-value follows the same recipe:

  1. Choose a test statistic that measures how far the data are from what H₀ predicts (z, t, χ², and so on).
  2. Calculate its value from your data.
  3. Find the probability, under H₀, of a test statistic at least that extreme. That probability is the p-value.

Worked z-test

A machine is supposed to fill bottles with 500 mL, and the fill volume has a known standard deviation of 4 mL. A sample of 25 bottles averages 498.2 mL. Is the machine underfilling?

z = (x̄ − μ0) ÷ (σ ÷ √n)

Standard error = 4 ÷ √25 = 0.8 mL

z = (498.2 − 500) ÷ 0.8 = −2.25

Two-tailed p = 2 × P(Z ≤ −2.25) ≈ 2 × 0.0122 = 0.024

If the machine were filling correctly, a sample mean this far from 500 mL in either direction would happen only about 2.4% of the time. At the 0.05 level, that is statistically significant evidence that the average fill has drifted.

If you had decided before collecting data to test only for underfilling, the one-tailed p-value would be 0.012. The p-value calculator converts any z, t, chi-square or F statistic into one- or two-tailed p-values, and how to calculate a z-score covers the z-statistic in depth.

Common z values and their p-values

|z| Two-tailed p One-tailed p
1.000 0.317 0.159
1.645 0.100 0.050
1.960 0.050 0.025
2.576 0.010 0.005
3.291 0.001 0.0005

For small samples with an estimated standard deviation, use the t-distribution instead; the t-test calculator handles one-sample, paired and two-sample tests.

How to interpret a p-value

Researchers compare the p-value with a significance level (α) chosen in advance, most often 0.05.

  • p ≤ α: reject H₀. The result is called statistically significant.
  • p > α: fail to reject H₀. The data do not provide strong enough evidence against it.

“Fail to reject” is deliberately different from “accept.” A non-significant result might mean there is no effect, or that the sample was too small to detect one.

The 0.05 threshold is a convention from the 1920s, associated with the statistician Ronald Fisher. Fields that run many tests or need stronger evidence use smaller thresholds; particle physicists require about five standard deviations (p ≈ 3 × 10⁻⁷ one-tailed) before announcing a discovery.

What a p-value is not

In 2016 the American Statistical Association issued a formal statement on p-values because misinterpretations had become so widespread. Its key points translate into these rules:

Common claim Why it is wrong
“p = 0.03 means a 3% chance the null is true.” The p-value assumes the null is true; it cannot also measure the probability that it is.
“p = 0.03 means a 97% chance the effect is real.” Same error in reverse.
“p > 0.05 proves there is no effect.” Absence of evidence is not evidence of absence.
“A smaller p means a bigger effect.” p depends on sample size as much as effect size.
“p = 0.049 and p = 0.051 are fundamentally different.” They are nearly identical evidence; the cutoff is arbitrary.

The ASA also stressed that scientific conclusions should not rest on whether a p-value crosses a single threshold, and that full reporting and transparency are essential.

Statistical vs practical significance

Two groups of 10,000 students take a test with a standard deviation of 15 points. One group averages 0.5 points higher.

Standard error of the difference = 15 × √(2 ÷ 10,000) ≈ 0.212

z = 0.5 ÷ 0.212 ≈ 2.36 → p ≈ 0.018

The difference is statistically significant, yet half a point on a 100-point test is meaningless in practice. The standardized effect size is 0.5 ÷ 15 ≈ 0.03 standard deviations, which is tiny. With enough data, almost any nonzero difference becomes significant, so always report the size of the effect alongside the p-value, ideally as a confidence interval. See how to calculate a confidence interval.

The multiple-testing trap

Each test at α = 0.05 has a 5% false-positive rate when the null is true. Run many tests and false positives pile up:

P(at least one false positive) = 1 − (1 − α)k

With 20 independent tests, that is 1 − 0.95²⁰ ≈ 64%. Testing until something turns up significant, sometimes called p-hacking, reliably produces findings that fail to replicate. The simplest fix, the Bonferroni correction, divides α by the number of tests: 0.05 ÷ 20 = 0.0025 per test.

Quick reference

  • The p-value is computed assuming H₀ is true.
  • Small p = data are surprising under H₀ = evidence against H₀.
  • Choose α and one- or two-tailed before looking at the data.
  • Report effect sizes and confidence intervals, not just p.

For tests built on the normal curve, the normal distribution calculator and the z-score calculator give the tail areas behind every p-value in this guide.

Frequently asked questions

What does a p-value of 0.05 mean?

It means that if the null hypothesis were true, results at least as extreme as yours would occur about 5% of the time, or 1 time in 20. It does not mean there is a 5% chance the null hypothesis is true or a 95% chance your finding is real.

Is a smaller p-value better?

A smaller p-value is stronger evidence against the null hypothesis, but it says nothing about how large or important the effect is. With a huge sample, a trivial difference can produce a tiny p-value. Always report the effect size and a confidence interval too.

What is the difference between a one-tailed and two-tailed p-value?

A two-tailed p-value counts extreme results in both directions, while a one-tailed p-value counts only the direction you predicted in advance. For a symmetric test, the one-tailed value is half the two-tailed value. Two-tailed tests are the default unless a direction was specified before seeing the data.

Why is 0.05 the cutoff for significance?

It is a convention popularized by the statistician Ronald Fisher in the 1920s, not a law of nature. Many fields use 0.01 or stricter thresholds, particle physics uses roughly 0.0000003 (five sigma), and the American Statistical Association advises against treating any single cutoff as a bright line.

Can a p-value prove the null hypothesis is true?

No. A large p-value only means the data are consistent with the null hypothesis; it does not show the null is correct. The study may simply have been too small to detect a real effect.