The chi-square statistic:
comparing what you saw with what you expected
Work through the sum, find the degrees of freedom, build expected counts for a two-way table, and know when the test is not valid.
Calcylator Editorial Team
Updated · 4 min read
What the statistic compares
The chi-square statistic, written χ², measures the gap between counts you observed and counts you would expect if a particular assumption were true. If the assumption is that a coin lands heads and tails equally, then 100 flips should give 50 and 50. Getting 30 heads and 70 tails is a gap, and χ² gives that gap a single number.
It works with counts of categories, not with averages. Each category contributes its own squared gap, scaled by what was expected there, and the contributions are added. Bigger gaps and smaller expected counts both push the statistic up.
- O:
- observed count in a category
- E:
- expected count in that category if the assumption holds
- Σ:
- add across all categories
The worked sum for the 30 and 70 case
Observed
Heads 30, tails 70
Expected
50 and 50
Heads
(30 − 50)² ÷ 50 = 400 ÷ 50 = 8
Tails
(70 − 50)² ÷ 50 = 400 ÷ 50 = 8
χ²
8 + 8 = 16
One degree of freedom, since two categories are fixed by their total of 100.
Squaring does two jobs. It removes the sign, so being 20 under in one category does not cancel being 20 over in another, and it makes large gaps count disproportionately. Dividing by E stops a gap of 20 on an expectation of 50 looking the same as 20 on an expectation of 5000.
Reading the number: degrees of freedom and critical values
A χ² of 16 means little until you compare it with the value that chance alone would exceed only rarely. That comparison depends on the degrees of freedom (df), which for a goodness-of-fit test is the number of categories minus one. Here it is 2 − 1 = 1.
| df | 5% critical value | 1% critical value |
|---|---|---|
| 1 | 3.841 | 6.635 |
| 2 | 5.991 | 9.210 |
| 3 | 7.815 | 11.345 |
| 5 | 11.070 | 15.086 |
With 1 df, 16 is far beyond 3.841 and also beyond 6.635. The p-value, the chance of a χ² this large or larger if the coin were fair, is about 0.00006, so the 30/70 split is strong evidence against a fair coin. A different, smaller χ², such as 2, would sit below 3.841 and be consistent with chance variation.
A two-way table: expected counts from the margins
For a table that cross-classifies two variables, the expected count in each cell comes from the row and column totals: row total × column total ÷ grand total. Suppose a store shows two versions of a page, each to 100 visitors, and 30 visitors buy after version A and 45 after version B.
| Bought | Did not buy | Row total | |
|---|---|---|---|
| Version A | 30 | 70 | 100 |
| Version B | 45 | 55 | 100 |
| Column total | 75 | 125 | 200 |
Expected, A bought
100 × 75 ÷ 200 = 37.5
Expected, A did not
100 × 125 ÷ 200 = 62.5
Expected, B bought
37.5
Expected, B did not
62.5
Cell contributions
1.5 + 0.9 + 1.5 + 0.9
χ²
4.8, with df = (2 − 1) × (2 − 1) = 1
Above 3.841, so the difference is significant at 5% (p ≈ 0.028).
A multi-category case: is a die fair?
Roll a die 60 times and record how often each face appears. A fair die gives an expected count of 10 for each of the six faces. Suppose you observe 8, 12, 9, 11, 10 and 10.
- (8 − 10)² ÷ 10 = 0.4
- (12 − 10)² ÷ 10 = 0.4
- (9 − 10)² ÷ 10 = 0.1
- (11 − 10)² ÷ 10 = 0.1
- 10 and 10 contribute 0 each
- Total χ² = 1.0 with 5 df
The 5% critical value for 5 df is 11.07, so a χ² of 1.0 is entirely consistent with a fair die. Note that very small values are also informative. An unnaturally small χ² can suggest figures that were smoothed or invented, since real randomness always produces some spread.
Statistical significance is not the size of the effect
A χ² value, even one that passes the critical value, says only that the gap is unlikely to be chance. It does not tell you whether the gap is large enough to matter. Effect-size measures put the statistic back on a scale of zero to one. Cramér's V is √(χ² ÷ (n × k)), where n is the number of observations and k is the smaller of rows − 1 and columns − 1.
- Two-version page test: V = √(4.8 ÷ (200 × 1)) = 0.155, a modest effect.
- Coin example: V = √(16 ÷ (100 × 1)) = 0.40, a large effect.
- A very large survey might give a significant χ² with V of 0.03, a difference that is real but of little practical weight.
There is also a link to the z-test for a proportion. With two categories, χ² equals the square of the z-score. For the coin, z = (30 − 50) ÷ √(100 × 0.5 × 0.5) = −4, and (−4)² = 16, which is exactly the χ² you computed. The two tests give identical p-values in that case.
When the approximation is unreliable
- Small expected counts: a common rule is that every expected count should be at least 5, or at least 80% of cells should be 5 or more.
- Dependent observations: the same person counted twice breaks the assumption of independence.
- Not counts: percentages and averages must be turned back into counts first.
- Tiny 2 × 2 tables: Fisher's exact test is preferred when expected counts are small.
- Huge samples: with enough data, trivial differences become significant, so look at the size of the effect too.
One more habit helps: write the hypothesis in words before computing anything. 'The die is fair' and 'version B converts no differently from version A' are testable statements, and they tell you which expected counts to build. Decide in advance which significance level you will use, commonly 5%, rather than choosing it after you see the p-value, since picking the threshold afterwards quietly turns a test into a justification.
A chi-square statistic calculator performs the arithmetic and looks up the p-value, but it cannot check these conditions for you. State the assumption being tested, confirm the counts are independent, and report both the statistic and the degrees of freedom. For anything with real consequences, have a statistician look over the design.
Common questions
What is the formula for the chi-square statistic?
χ² = Σ (O − E)² ÷ E, where O is the observed count and E is the expected count in each category. For observed 30 and 70 against expected 50 and 50, it is 8 + 8 = 16.
How do I find the degrees of freedom?
For a goodness-of-fit test, use the number of categories minus one. For a contingency table, use (rows − 1) × (columns − 1). A 2 × 2 table therefore has 1 degree of freedom, and a six-sided die test has 5.
What does a chi-square of 16 mean?
With 1 degree of freedom, 16 is much larger than the 5% critical value of 3.841 and the 1% value of 6.635. The p-value is about 0.00006, so the observed counts are very unlikely if the expected split were true.
How do I calculate expected counts in a table?
Multiply the row total by the column total and divide by the grand total. If a row of 100 and a column of 75 sit in a table of 200, the expected count in that cell is 100 × 75 ÷ 200 = 37.5.
When should I not use chi-square?
Avoid it when expected counts are below about 5, when observations are not independent, or when your data are percentages or measurements rather than counts. For small 2 × 2 tables, Fisher's exact test is generally preferred.
Was this guide helpful?
Continue reading
View all blogsT-Test Statistic Formula With a Worked Example
The one-sample t statistic is (mean − claimed mean) ÷ (SD ÷ √n). With mean 54, claim 50, SD 10 and n = 25, t = 2.0 on 24 degrees of freedom.
5 min read
Z-Score Formula: Standard Deviations From the Mean
A z-score subtracts the mean and divides by the standard deviation: (85 − 70) ÷ 10 = 1.5. That puts the value about the 93rd percentile if scores are normal.
5 min read
Variance and Standard Deviation: n vs n−1
Population variance divides by n, sample variance by n − 1. See both on the same six numbers: 2.92 versus 3.5, and SD 1.71 versus 1.87.
6 min read




