Chi-Square Test Calculator

Test whether counts match an expected distribution, or whether two categorical variables are related, with expected counts, cell contributions and the p-value.

Test
One count per category, separated by commas or spaces.
p-value
0.41588
Degrees of freedom
5
Critical χ² (α = 0.05)
11.0705
Decision
Fail to reject H₀
Categories
6
Total count (N)
120
Chi-square statistic (χ²)5df = 5, p = 0.41588 → Fail to reject H₀
  • Assumes independent observations, each falling in exactly one category.

Show the work

  1. Equal expected counts: N ÷ k = 120 ÷ 6 = 20 per category
  2. χ² = Σ (O − E)² ÷ E = (25 − 20)²/20 + (17 − 20)²/20 + (15 − 20)²/20 + (23 − 20)²/20 + (24 − 20)²/20 + (16 − 20)²/20 = 5
  3. Degrees of freedom: k − 1 = 6 − 1 = 5
  4. p-value = P(χ²5 ≥ 5) = 0.41588; critical value at α = 0.05 is 11.0705
  5. p = 0.41588 > α = 0.05, so the result is not statistically significant at the 5% level: do not reject the null hypothesis.
05101520χ² = 5

Blue: p-value area beyond your χ². Amber: rejection region beyond the critical value 11.0705.

Observed vs. expected
CategoryObservedExpectedO − E(O − E)² ÷ E
1252051.25
21720−30.45
31520−51.25
4232030.45
5242040.8
61620−40.8
Total12012005

The chi-square test is the standard tool for counts. It asks whether the numbers you observed in each category are close enough to what a hypothesis predicts that the gaps could be chance. The calculator runs both versions: goodness of fit for a single list of counts, and independence for a two-way contingency table. It shows the expected counts, how much each cell contributes to the statistic, the p-value and the critical value on a shaded chi-square curve.

How to use the chi-square calculator

  1. Choose Goodness of fit or Independence.
  2. For goodness of fit, enter the observed count in each category. Pick “equally likely” or enter the expected proportions, percentages, a ratio such as 9:3:3:1, or expected counts. Category names are optional.
  3. For independence, type the table one row per line. You can start each line with a label and a colon, and give column names separated by commas.
  4. Set the significance level α and read the decision, then check the notes for small expected counts.

Chi-square formula

Both tests use the same statistic, summed over every category or cell:

χ² = Σ (O − E)² ÷ E

For goodness of fit, E is the total count times each category’s expected share, and df = k − 1 for k categories. For a test of independence, the expected count in row i, column j is

Eij = (row totali × column totalj) ÷ N

with df = (rows − 1)(columns − 1). The p-value is the area to the right of χ² under the chi-square distribution with those degrees of freedom. Effect size for a table is Cramér’s V = √(χ² ÷ (N × (min(r, c) − 1))).

Worked examples

Is the die fair? A die rolled 120 times gives faces 1–6 with counts 25, 17, 15, 23, 24 and 16. If fair, each face is expected 20 times.

χ² = (5² + 3² + 5² + 3² + 4² + 4²) ÷ 20 = 100 ÷ 20 = 5.0, with 5 degrees of freedom. The p-value is 0.4159, well above 0.05, and the critical value is 11.07. The counts are entirely consistent with a fair die.

Is outcome related to treatment? In a trial, 80 patients received a treatment and 80 a placebo:

Improved No change Worse
Treatment 42 31 7
Placebo 28 37 15

The expected count for Treatment/Improved is 80 × 70 ÷ 160 = 35. Summing all six contributions gives χ² = 6.2385 with 2 df and p = 0.0442, so at α = 0.05 the outcome distribution differs between groups. Cramér’s V is 0.1975, a weak-to-moderate association. The contribution table shows where the difference lies: more improvement and fewer worsened patients under treatment than expected.

Reading the result

Look at the cells, not just the p-value

A significant χ² says the pattern is not what independence predicts, but not which cells drive it. Compare observed with expected and look for the largest contributions. Cells with a large positive O − E are overrepresented.

Degrees of freedom when parameters are estimated

If you estimate parameters from the data to build the expected distribution (for example, fitting a Poisson mean before testing fit), subtract one degree of freedom per estimated parameter. This calculator assumes the expected distribution was fixed in advance; adjust df by hand and use the p-value calculator if needed.

Large samples find tiny effects

With thousands of observations, trivial departures become significant. Cramér’s V and the observed percentages tell you whether the association matters in practice. To build the counts in the first place, the frequency table calculator tallies raw categorical data.

Each observation must be independent and counted in exactly one category or cell. Repeated measurements on the same subjects need a different test, such as McNemar's for paired yes/no data.

Frequently asked questions

What is the difference between goodness of fit and independence?

Goodness of fit compares one categorical variable with a distribution you specify in advance, such as a fair die or a 9:3:3:1 genetic ratio. Independence uses a two-way table of two variables, such as treatment and outcome, and asks whether the row and column variables are related. The arithmetic is similar; the expected counts and degrees of freedom differ.

Why must expected counts be at least 5?

The chi-square p-value relies on an approximation that works well only when expected counts are not too small. A common rule is that all expected counts should be at least 1 and no more than 20% of them below 5. If your table breaks the rule, merge sparse categories or use Fisher's exact test for a 2 × 2 table.

Can I enter percentages instead of counts?

Not for the observed data. The test depends on how many observations you have, and percentages hide that. Expected values can be given as proportions, percentages or a ratio because the calculator rescales them to the observed total.

When should I use the Yates correction?

It is an optional adjustment for 2 × 2 tables that makes the test more conservative. Many statisticians now skip it because it overcorrects; with small counts Fisher's exact test is the better fix. The calculator applies it only when you tick the box and the table is 2 × 2.

What does Cramér's V measure?

It turns the χ² statistic into a 0-to-1 measure of how strongly two categorical variables are associated, independent of sample size. Roughly, 0.1 is weak, 0.3 moderate and 0.5 strong, though conventions vary by field and table size.

Last reviewed October 2026 by the CalcFluent editorial team. How we check our calculators.