Software and textbooks often report a test statistic and leave you to find the p-value, or you may have a statistic from a paper and want to check its significance. This calculator converts z, t, chi-square (χ²) and F statistics into exact left-tail, right-tail and two-tailed p-values, shows the critical value at your chosen α, and shades the p-value area under the matching curve.
How to use the p-value calculator
- Pick the distribution of your statistic: z, t, χ² or F.
- Enter the statistic and its degrees of freedom (one df for t and χ², two for F).
- Choose the tail. Two-tailed is standard for z and t tests; right-tailed is standard for χ² and F tests.
- Set α to see the decision and critical value. The tape also lists both tail areas so you can switch tails without retyping.
How p-values are computed
Each p-value is a tail area under a probability curve:
The z tail uses the standard normal distribution. For t, χ² and F the calculator evaluates the exact cumulative distribution through the regularized incomplete beta and gamma functions, for example
These routines agree with printed tables to every digit shown (for instance t₀.₀₂₅,₁₀ = 2.228, χ²₀.₀₅,₁ = 3.841 and F₀.₀₅;₂,₁₀ = 4.103) and keep their precision far into the tails.
Worked examples
t statistic. A one-sample t-test with 16 observations gives t = 2.1, so df = 15. The right-tail area is 0.026528. Doubling it gives a two-tailed p = 0.0531, just above 0.05, and the two-tailed critical value is ±2.1314. At α = 0.05 you would not reject H₀, though the result is close.
z statistic. z = 1.96 gives a two-tailed p of 0.0500, the familiar boundary for 95% confidence.
Chi-square. A test of independence on a 2 × 3 table produced χ² = 6.2385 with 2 df. The right-tail p-value is 0.0442, below 0.05; the critical value is 5.9915.
F. An ANOVA comparing 4 groups with 24 observations has df = 3 and 20. An F of 4.1 gives p = 0.0202, beyond the 5% critical value of 3.0984.
Interpreting p-values responsibly
What a p-value is not
- It is not the probability that H₀ is true.
- It is not the probability that the result happened “by chance.”
- It does not measure the size or importance of an effect. A huge study can make a trivial difference highly significant.
Common thresholds
| p-value | Typical wording |
|---|---|
| p ≤ 0.001 | very strong evidence against H₀ (***) |
| p ≤ 0.01 | strong evidence (**) |
| p ≤ 0.05 | moderate evidence (*) |
| 0.05 < p ≤ 0.10 | weak evidence, sometimes called marginal |
| p > 0.10 | little or no evidence |
Multiple comparisons
Each test at α = 0.05 has a 5% false-positive rate, so running 20 tests on pure noise produces about one “significant” result. When testing many hypotheses, adjust α (for example, Bonferroni: divide α by the number of tests) or use a method built for the purpose.
To compute the statistic from raw data first, use the t-test calculator, the chi-square calculator or the ANOVA calculator. For critical values on their own, the t-distribution calculator prints a full table.
The p-value assumes the test's own conditions hold (independence, approximate normality where required, adequate expected counts). A precise p-value from a misapplied test is still misleading.
Frequently asked questions
What is a p-value in plain terms?
It is the probability of getting a test statistic at least as extreme as yours if the null hypothesis were exactly true. A p-value of 0.03 means that, in a world with no real effect, results this far out would occur about 3% of the time. Small values make the null hypothesis look less believable.
Is p = 0.05 a magic threshold?
No. The 0.05 cutoff is a convention, not a law of nature. A p-value of 0.049 and one of 0.051 carry nearly the same evidence. Report the exact p-value, the effect size and a confidence interval rather than only significant or not significant.
When should I use a one-tailed p-value?
Only when the research question is directional and was fixed before you saw the data, for example testing whether a new process is faster, where a slower result would be treated the same as no change. Otherwise use the two-tailed value, which is the default in most software and journals.
Why are chi-square and F tests usually right-tailed?
Both statistics grow as the data move away from the null hypothesis, and they cannot be negative. Large values are the evidence against H₀, so the p-value is the right-tail area. Left-tailed and two-tailed versions exist for special cases, such as testing whether a variance is suspiciously small.
Why does my t p-value differ from a z p-value for the same number?
The t distribution has heavier tails than the normal, especially with few degrees of freedom, so the same statistic is less extreme under t. With df = 15, t = 2.1 gives a two-tailed p of about 0.053, while z = 2.1 gives about 0.036. As df grows, the two converge.