The t-test is the workhorse for comparing means when the population standard deviation is unknown, which is nearly always. This calculator runs all four common versions, accepts either raw data or summary statistics, and reports the t statistic, degrees of freedom, p-value, critical value, a confidence interval for the difference and Cohen’s d, with a shaded t curve showing where your result falls.
How to use the t-test calculator
- Pick the test: one-sample (a mean against a target value), two-sample Welch or pooled (two independent groups), or paired (matched measurements).
- Choose raw data to paste values, or summary statistics to type means, standard deviations and sample sizes.
- Enter the hypothesized mean or difference. For group comparisons this is usually 0.
- Choose the alternative hypothesis: two-tailed (≠), left-tailed (<) or right-tailed (>). Decide this before looking at the data.
- Set α, commonly 0.05. Read the verdict, then check the summary table and the notes on assumptions.
T-test formulas
One-sample and paired tests use the same formula; for a paired test the data are the differences d within each pair:
Welch’s two-sample test:
with Welch–Satterthwaite degrees of freedom. The pooled test replaces the denominator with sp√(1/n₁ + 1/n₂), where sp² is the weighted average of the two variances, and uses df = n₁ + n₂ − 2. The p-value is the tail area of the t distribution beyond the observed t.
Worked example
A gardener compares plant heights (cm) under two fertilizers, 10 plants each.
- Fertilizer A: 23, 25, 28, 30, 26, 27, 24, 29, 31, 26 → mean 26.9, SD 2.6013
- Fertilizer B: 20, 22, 25, 21, 24, 23, 19, 26, 22, 21 → mean 22.3, SD 2.2136
Welch’s test, two-tailed, α = 0.05:
- Standard error: √(2.6013²/10 + 2.2136²/10) = √(0.676667 + 0.49) = 1.080123.
- t = (26.9 − 22.3) ÷ 1.080123 = 4.2588.
- Degrees of freedom: 1.166667² ÷ (0.676667²/9 + 0.49²/9) = 17.55.
- Two-tailed p-value: 0.000497, far below 0.05, so reject H₀.
- 95% confidence interval for the difference: 4.6 ± 2.1048 × 1.080123 = 2.33 to 6.87 cm.
- Cohen’s d = 4.6 ÷ 2.4152 (pooled SD) = 1.90, a large effect.
The pooled test gives nearly the same answer here (df = 18, p = 0.000472) because the two SDs are similar. If the same numbers had been before-and-after measurements on 10 plants, the paired test would give t = 5.81 with 9 df.
Interpreting the result
Significance is not size
A tiny difference can be statistically significant in a huge sample, and a large one can miss significance in a small sample. Always read the confidence interval: here the data are consistent with a true difference anywhere from about 2.3 to 6.9 cm. If that whole range is practically meaningful, the result is useful; if it straddles values that matter and values that do not, more data are needed.
One-tailed tests
A one-tailed test halves the p-value when the effect goes in the predicted direction, which is why it must be chosen in advance. Switching to one-tailed after seeing the data inflates the false-positive rate.
More than two groups
Running many pairwise t-tests multiplies the chance of a false positive. For three or more groups, start with the one-way ANOVA calculator. To look up a t critical value directly, use the t-distribution calculator, and to convert any t statistic into a p-value use the p-value calculator.
Results assume independent observations. For heavily skewed small samples or ordinal data, consider a rank-based test such as Mann–Whitney or Wilcoxon signed-rank.
Frequently asked questions
Should I use Welch's t-test or the pooled t-test?
Use Welch's test by default. It does not assume the two groups share a variance, loses almost nothing when they do, and protects you when they do not. The pooled test is the classic textbook version and is fine when the standard deviations are similar and the group sizes are close.
When is a paired t-test the right choice?
When each value in one sample is naturally matched to one value in the other: the same person before and after, two measurements on the same item, or twins assigned to different treatments. Pairing removes the person-to-person variation and usually gives a much more sensitive test.
What does the p-value tell me?
It is the probability of seeing a t statistic at least as extreme as yours if the null hypothesis were true. A small p-value says the data would be surprising under the null. It does not give the probability that the null is true, and it says nothing about how large or important the difference is.
Why are the degrees of freedom not a whole number?
Welch's test estimates its degrees of freedom with the Welch–Satterthwaite formula, which blends the two sample variances and sizes. A fractional value such as 17.55 is normal; the p-value is computed from the t distribution with exactly that df.
How do I interpret Cohen's d?
Cohen's d expresses the difference in standard-deviation units. Rough conventions are 0.2 small, 0.5 medium and 0.8 large, but what counts as meaningful depends on the field. Report it alongside the p-value so readers can judge practical importance.
What assumptions does the t-test make?
Independent observations, and either roughly normal populations or samples large enough (around 30 per group) for the central limit theorem to help. The test is fairly robust to mild skew but sensitive to extreme outliers, especially in small samples.