Chi-Squared Test in Psychology: Formula, Example, Table

For the first twenty-two years of its life, the chi-squared test used the wrong degrees of freedom. Karl Pearson gave a 2×2 table 3 instead of 1, until Ronald Fisher corrected him in 1922 (Stigler, 2008).
Key Takeaways
- Use a chi-squared test when you are testing for a difference or an association, your data are nominal frequencies in named categories, and each participant appears in one cell only.
- The formula is chi-squared equals the sum of (observed minus expected) squared, divided by expected. The expected frequency for any cell is its row total times its column total, divided by the overall total.
- Degrees of freedom are (rows minus 1) times (columns minus 1), and your calculated value must equal or exceed the critical value in the table for the result to be significant.
The chi-squared test is the one inferential test in A level psychology that works on counts rather than scores. Every other test on the specification wants numbers that measure something: seconds recalled, words remembered, a rating out of ten. Chi-squared wants only tallies. How many children chose the construction area, how many chose the role-play corner, and whether that split depends on something else.
That makes it both the easiest test to spot and the easiest to get wrong. Spotting it is straightforward once you know that nominal data and an independent groups design point almost exclusively at chi-squared, which is the reasoning set out in our guide to choosing the right statistical test. Getting it wrong usually comes from one of three places: calculating expected frequencies from the wrong totals, misreading the degrees of freedom, or picking the wrong column in the critical values table because nobody explained what one-tailed means for a distribution that only has one tail.
This guide works through all of it. There is a full worked example with numbers that come out cleanly, a second example with a larger table so you can see what changes, a free calculator, and the complete critical values table exactly as the exam boards print it. If you are unsure whether your data even are nominal, start with levels of measurement and come back.
What is the chi-squared test in psychology?
The chi-squared test is an inferential statistical test that compares the frequencies you actually observed with the frequencies you would expect if there were no difference or no association between your variables. If the gap between observed and expected is large enough to be unlikely by chance, the result is significant.
The symbol is the Greek letter chi, written as a lower-case x and squared: it is pronounced “kye”, to rhyme with “sky”, not “chee” or “chai”. Karl Pearson introduced the test in 1900 in a paper with one of the longer titles in the history of statistics, and it has been part of the standard toolkit ever since (Pearson, 1900). It is sometimes called Pearson’s chi-squared test to distinguish it from later variants.
The logic is worth pausing on, because it is the same logic that sits under every inferential test. You start by assuming nothing is going on: the null hypothesis. Under that assumption you can work out what the data ought to look like. Then you measure how far the real data have drifted from that expectation. Chi-squared is simply a number that summarises the total drift across every cell of the table. A small chi-squared means the data look roughly like the null hypothesis predicted. A large one means they do not.
What makes the test distinctive is the kind of data it accepts. It does not care how much of something a participant produced. It cares only which box they went into. That is why it is the natural partner for observational research, content analysis and any study that codes behaviour into named categories rather than measuring it on a scale.
When should you use a chi-squared test?
Use a chi-squared test when all three of the following are true: you are testing for a difference or an association rather than a correlation between two measured scores, your data are nominal frequencies, and your design is independent groups so that each participant contributes to exactly one cell.
The three conditions in full
Nominal data. Your results must be counts of how many participants fall into each named category. “Twenty-eight children helped, twelve did not” is nominal. “Children took an average of 4.2 seconds to help” is not: that is interval data and needs a t-test. The categories must be genuinely separate, with no overlap and no ordering that matters.
Independent data. Each participant must appear once and once only, in a single cell. This is the condition students most often break. If you observe the same twenty children twice, before and after an intervention, and then build a table of forty tallies, the test is invalid: twenty of those tallies are not independent of the other twenty. For repeated measures on nominal data you would need a different test, such as the sign test.
A test of difference or association. Chi-squared answers questions of the form “does the split across these categories depend on which group you are in?” It cannot tell you the direction or the size of an effect on its own, and it cannot handle two continuous variables, which is the job of a correlation coefficient.
Is chi-squared a test of difference or a test of association?
It is both, and the exam boards genuinely disagree about how to label it. Pearson Edexcel lists it in its specification as “chi squared (for difference) tests”, while the same document elsewhere describes it as a test of association between categories (Pearson Edexcel, 2026). AQA simply names it among the tests students must know when to use (AQA, 2015).
The disagreement is cosmetic. Mathematically it is one test. Whether you describe the outcome as “boys and girls differed in which area they chose” or “there was an association between sex and area chosen” depends on how you framed the hypothesis, not on anything in the arithmetic. Write your hypothesis using whichever wording your board prefers, and be consistent between the hypothesis and the conclusion.
The chi-squared formula explained
The chi-squared formula is the sum, across every cell in the table, of the observed frequency minus the expected frequency, squared, then divided by the expected frequency.
χ2 = Σ ( O − E )2 ÷ E
Σ means “add up across all cells” · O is the observed frequency · E is the expected frequency
Each part of that formula is doing a specific job, and understanding why makes the whole test far easier to remember.
Why subtract? O minus E is the raw discrepancy for one cell: how far reality drifted from the null hypothesis in that particular box.
Why square? Because the discrepancies always add to zero across the whole table. Every cell that comes out above expectation is balanced by one that comes out below. Squaring removes the signs so that the discrepancies accumulate instead of cancelling, and it also makes large discrepancies count for disproportionately more than small ones.
Why divide by E? This is the step that gives the test its sense of proportion. A discrepancy of 10 is enormous if you only expected 12 people in that cell and trivial if you expected 500. Dividing by the expected frequency scales each discrepancy against how big it should have been, so every cell contributes fairly regardless of how many people passed through it.
The second formula you need is the one for expected frequencies, and it is not in the main formula above because it has to be worked out separately for every cell before you can start.
E = ( row total × column total ) ÷ overall total
That expression is the null hypothesis written as arithmetic. It says: if the two variables really had nothing to do with each other, then each row would split across the columns in exactly the same proportions as the table does overall. Everything else in the test is a measurement of how badly that assumption fails.
Chi-squared calculator and critical values table
Enter your observed frequencies below and the calculator works out the expected frequencies, the chi-squared value, the degrees of freedom and the verdict against the critical value. The full critical values table sits underneath it, and is the same table the exam boards print in their formula booklets.
Chi-Squared Test Calculator
Choose the size of your contingency table, type in the observed frequencies, and press calculate. Nominal data and independent groups only.
Critical values of chi-squared. In each column the upper figure is the significance level for a one-tailed (directional) test and the lower figure is the level for a two-tailed (non-directional) test. Your calculated value must equal or exceed the critical value for the result to be significant.
| df | Level of significance | |||||
|---|---|---|---|---|---|---|
| 0.10 0.20 | 0.05 0.10 | 0.025 0.05 | 0.01 0.025 | 0.005 0.01 | 0.0005 0.001 | |
| 1 | 1.64 | 2.71 | 3.84 | 5.02 | 6.64 | 10.83 |
| 2 | 3.22 | 4.61 | 5.99 | 7.38 | 9.21 | 13.82 |
| 3 | 4.64 | 6.25 | 7.82 | 9.35 | 11.35 | 16.27 |
| 4 | 5.99 | 7.78 | 9.49 | 11.14 | 13.28 | 18.47 |
| 5 | 7.29 | 9.24 | 11.07 | 12.83 | 15.09 | 20.52 |
| 6 | 8.56 | 10.65 | 12.59 | 14.45 | 16.81 | 22.46 |
| 7 | 9.80 | 12.02 | 14.07 | 16.01 | 18.48 | 24.32 |
| 8 | 11.03 | 13.36 | 15.51 | 17.54 | 20.09 | 26.12 |
| 9 | 12.24 | 14.68 | 16.92 | 19.02 | 21.67 | 27.88 |
| 10 | 13.44 | 15.99 | 18.31 | 20.48 | 23.21 | 29.59 |
| 11 | 14.63 | 17.28 | 19.68 | 21.92 | 24.73 | 31.26 |
| 12 | 15.81 | 18.55 | 21.03 | 23.34 | 26.22 | 32.91 |
| 13 | 16.99 | 19.81 | 22.36 | 24.74 | 27.69 | 34.53 |
| 14 | 18.15 | 21.06 | 23.69 | 26.12 | 29.14 | 36.12 |
| 15 | 19.31 | 22.31 | 25.00 | 27.49 | 30.58 | 37.70 |
| 16 | 20.47 | 23.54 | 26.30 | 28.85 | 32.00 | 39.25 |
| 17 | 21.62 | 24.77 | 27.59 | 30.19 | 33.41 | 40.79 |
| 18 | 22.76 | 25.99 | 28.87 | 31.53 | 34.81 | 42.31 |
| 19 | 23.90 | 27.20 | 30.14 | 32.85 | 36.19 | 43.82 |
| 20 | 25.04 | 28.41 | 31.41 | 34.17 | 37.57 | 45.32 |
| 25 | 30.68 | 34.38 | 37.65 | 40.65 | 44.31 | 52.62 |
| 30 | 36.25 | 40.26 | 43.77 | 46.98 | 50.89 | 59.70 |

Chi-squared worked example, step by step
The worked example below uses a hypothetical observational study with numbers chosen to come out cleanly, so you can follow every step without a calculator getting in the way. Once you have done it once with tidy numbers, messy ones hold no surprises.
The study. A researcher wants to know whether the physical setting affects how young children play. She observes 120 four-year-olds during free play, 60 of them outdoors and 60 indoors, and codes each child’s dominant play behaviour as either cooperative or solitary. Each child is observed once, in one setting only, and appears in exactly one cell.
The hypothesis. There will be a difference in the type of play observed between children playing outdoors and children playing indoors. This is non-directional, so the test is two-tailed.
Step 1: Set out the contingency table
Draw the table with the independent variable down the side and the categories of behaviour across the top, then add the row totals, the column totals and the overall total. Getting these margins right matters more than anything else in the test, because every expected frequency is built from them.
| Setting | Cooperative play | Solitary play | Row total |
|---|---|---|---|
| Outdoors | 45 | 15 | 60 |
| Indoors | 30 | 30 | 60 |
| Column total | 75 | 45 | 120 |
Step 2: Calculate the expected frequencies
The expected frequency for each cell is its row total multiplied by its column total, divided by the overall total. Do this for every cell, keeping at least two decimal places until the very end.
- Outdoors, cooperative: (60 × 75) ÷ 120 = 4500 ÷ 120 = 37.5
- Outdoors, solitary: (60 × 45) ÷ 120 = 2700 ÷ 120 = 22.5
- Indoors, cooperative: (60 × 75) ÷ 120 = 4500 ÷ 120 = 37.5
- Indoors, solitary: (60 × 45) ÷ 120 = 2700 ÷ 120 = 22.5
Two checks are worth doing before you go on. First, the expected frequencies in each row must add up to that row’s total: 37.5 plus 22.5 is 60, correct for both rows. Second, every expected frequency here is comfortably above 5, which is the assumption discussed later in this article. If either check fails, stop and find the arithmetic error before continuing.
Notice what these expected frequencies are saying. Overall, 75 of the 120 children played cooperatively, which is 62.5 per cent. If setting made no difference at all, you would expect 62.5 per cent of each group of 60 to play cooperatively, which is 37.5 children. That is all the formula is doing: applying the overall proportion to each row.
Step 3: Find O minus E for every cell
Subtract each expected frequency from its observed frequency. In a 2×2 table the same number appears in all four cells, with the signs alternating, which is a useful self-check.
- Outdoors, cooperative: 45 − 37.5 = +7.5
- Outdoors, solitary: 15 − 22.5 = −7.5
- Indoors, cooperative: 30 − 37.5 = −7.5
- Indoors, solitary: 30 − 22.5 = +7.5
These four numbers add to zero, which they always will. That is exactly why the next step squares them.
Step 4: Square each difference and divide by E
Square each difference, then divide by that cell’s expected frequency. Here 7.5 squared is 56.25 for every cell, so only the divisor changes.
| Cell | O | E | O − E | (O − E)² | (O − E)² ÷ E |
|---|---|---|---|---|---|
| Outdoors, cooperative | 45 | 37.5 | +7.5 | 56.25 | 1.50 |
| Outdoors, solitary | 15 | 22.5 | −7.5 | 56.25 | 2.50 |
| Indoors, cooperative | 30 | 37.5 | −7.5 | 56.25 | 1.50 |
| Indoors, solitary | 30 | 22.5 | +7.5 | 56.25 | 2.50 |
| Total | 120 | 120 | 0 | 8.00 |
Step 5: Add the cells to get chi-squared
Add the final column: 1.50 plus 2.50 plus 1.50 plus 2.50 gives a calculated chi-squared value of 8.00. This is the observed value, sometimes called the calculated value or the test statistic. On its own it means nothing at all until you compare it with a critical value.
Step 6: Work out the degrees of freedom
Degrees of freedom for a chi-squared test are (number of rows minus 1) multiplied by (number of columns minus 1). Count only the data rows and columns, never the totals.
This table has 2 data rows and 2 data columns, so df = (2 − 1) × (2 − 1) = 1 × 1 = 1. Every 2×2 contingency table has 1 degree of freedom, however many participants it contains.
Step 7: Compare with the critical value
Find the row for df = 1 in the critical values table, then the column for a two-tailed test at the 0.05 level of significance. The critical value is 3.84.
The calculated value of 8.00 is larger than 3.84, so the result is significant at p < 0.05. In fact it also clears the 0.01 critical value of 6.64, so it can be reported at the stricter level. It does not reach the 0.001 critical value of 10.83.
Note the direction of the comparison, because it is the opposite of the sign test. For chi-squared the calculated value must be equal to or greater than the critical value. For the sign test, Wilcoxon and Mann-Whitney it must be equal to or less than. Mixing these up turns a correct calculation into a wrong conclusion, and it is one of the most common ways marks are lost.
Step 8: State the conclusion
A conclusion needs three things: the decision about the null hypothesis, the significance level, and a plain statement of what actually happened in the data.
“The calculated value of chi-squared (8.00) exceeds the critical value (3.84) for df = 1 at p < 0.05 for a two-tailed test. The null hypothesis is therefore rejected. There is a significant difference in the type of play observed between the outdoor and indoor settings, with cooperative play more frequent outdoors than indoors.”
That last clause matters. Chi-squared itself is direction-blind: a value of 8.00 tells you the table departs from independence but not which way. You get the direction by looking back at the original frequencies and describing what you see. An answer that stops at “the result was significant” has not finished the job.
A second worked example: a 2×3 table
Larger tables work in exactly the same way, but two things change: the expected frequencies stop being symmetrical, and the degrees of freedom rise. This example uses a 2×3 table, which is the shape you get whenever you compare two groups across three categories.
The study. A hypothetical researcher assesses 100 infants in each of two countries using the Strange Situation and records each infant’s attachment classification as secure, insecure-avoidant or insecure-resistant. This is the same design as the real cross-cultural work on attachment, and if you want the genuine findings rather than invented ones, van IJzendoorn and Kroonenberg’s (1988) meta-analysis of 32 samples from 8 countries is the study to read. Our article on secure attachment covers what the classifications mean.
| Country | Secure | Avoidant | Resistant | Row total |
|---|---|---|---|---|
| Country A | 60 | 25 | 15 | 100 |
| Country B | 50 | 10 | 40 | 100 |
| Column total | 110 | 35 | 55 | 200 |
The expected frequencies use the same formula. For Country A, secure: (100 × 110) ÷ 200 = 55. For Country A, avoidant: (100 × 35) ÷ 200 = 17.5. For Country A, resistant: (100 × 55) ÷ 200 = 27.5. Because the two row totals happen to be equal here, Country B has the same three expected frequencies.
| Cell | O | E | O − E | (O − E)² ÷ E |
|---|---|---|---|---|
| A, secure | 60 | 55.0 | +5.0 | 0.45 |
| A, avoidant | 25 | 17.5 | +7.5 | 3.21 |
| A, resistant | 15 | 27.5 | −12.5 | 5.68 |
| B, secure | 50 | 55.0 | −5.0 | 0.45 |
| B, avoidant | 10 | 17.5 | −7.5 | 3.21 |
| B, resistant | 40 | 27.5 | +12.5 | 5.68 |
| Total | 200 | 200 | 0 | 18.70 |
Degrees of freedom are (2 − 1) × (3 − 1) = 1 × 2 = 2. The critical value at df = 2 for a two-tailed test at p < 0.05 is 5.99, at p < 0.01 it is 9.21, and at p < 0.001 it is 13.82. The calculated value of 18.70 exceeds all three, so the null hypothesis is rejected at p < 0.001.
The cell-by-cell contributions are worth reading even though no exam asks you to. Chi-squared gives one number for the whole table, but that number is built from parts, and the parts tell you the story. Here the secure cells contribute 0.45 each and the resistant cells 5.68 each. The countries are not really differing in security at all; they are differing in which kind of insecurity is more common. A blanket statement that “attachment type differed between countries” would be true but would miss the finding entirely, a point our article on cultural bias in psychology returns to in a wider context.
What a real exam question looks like
Exam questions rarely ask for the whole calculation from scratch. They give you part of it and test whether you understand the rest. OCR’s sample assessment material for H567/01 Research Methods is a good illustration of the pattern, and it is worth walking through because the same three questions appear in various forms across all the boards (OCR, 2026).
The scenario involves participants shown either animals or kitchen items, then asked what they perceived in an ambiguous image. The results form a 2×2 table.
| Items presented | Perceived as monkey | Perceived as teapot |
|---|---|---|
| Animals | 15 | 10 |
| Kitchen items | 5 | 12 |
Question one asks for two reasons why chi-squared was the right test. The mark scheme awards credit for saying that the study investigated a difference or association, that the design was unrelated, and that the data were nominal, with further marks for tying each reason to the study itself. These are the three conditions set out earlier in this article, which is why they are worth learning as a set rather than individually.
Question two asks how the degrees of freedom would be determined. Two data rows and two data columns give (2 − 1) × (2 − 1) = 1.
Question three gives the calculated value as 3.80 and a short table of one-tailed critical values — 2.71 at the 0.05 level, 3.84 at 0.025 and 5.41 at 0.01 — and asks whether the result is significant. The calculated value of 3.80 exceeds 2.71, so the result is significant at the 0.05 level. It falls just short of 3.84, so it cannot be claimed at 0.025. A complete answer says both things.
It is worth checking the given value yourself, because doing so proves the whole method. The row totals are 25 and 17, the column totals 20 and 22, and the overall total 42. The expected frequencies are therefore 11.90, 13.10, 8.10 and 8.90. Each observed frequency differs from its expected frequency by 3.10, and the four contributions come to 0.80, 0.73, 1.18 and 1.08, which add to 3.80. The exam board’s figure and the method in this article agree exactly.
Two things are worth noticing about this example. First, one expected frequency is 8.10 and all four clear 5, so the assumption holds; with a total of only 42 participants it easily might not have. Second, this is a case where the choice of column genuinely decides the answer. Had the hypothesis been non-directional, 3.80 would have been compared against 3.84 and the result would have been non-significant. The difference between a pass and a fail on that question is understanding which column to read.
Degrees of freedom in the chi-squared test
Degrees of freedom are the number of cells in a contingency table that could vary freely once the row and column totals are fixed. For a chi-squared test the formula is df = (rows − 1) × (columns − 1), counting only the data rows and columns.
The idea behind the formula is easier to see than to define. Take the 2×2 table from the worked example and imagine the four margins are already known: 60 outdoors, 60 indoors, 75 cooperative, 45 solitary. Now try to fill in the cells. Put any number you like in the first cell, say 45. Every other cell is then forced: the rest of that row must be 15 to reach 60, the rest of that column must be 30 to reach 75, and the final cell must be 30. One cell was free, and one is the degrees of freedom.
Do the same with a 2×3 table and you find two cells are free before the rest are forced. A 3×4 table gives six. The pattern is always (rows − 1) × (columns − 1).
| Table size | Calculation | Degrees of freedom |
|---|---|---|
| 2 x 2 | 1 × 1 | 1 |
| 2 x 3 | 1 × 2 | 2 |
| 2 x 4 | 1 × 3 | 3 |
| 3 x 3 | 2 × 2 | 4 |
| 3 x 4 | 2 × 3 | 6 |
| 4 x 5 | 3 × 4 | 12 |
This formula is precisely what Karl Pearson got wrong. His 1900 paper treated the degrees of freedom as the number of cells minus one, which for a 2×2 table gives 3 rather than 1, and for a 3×3 table gives 8 rather than 4 (Stigler, 2008). The error survived for two decades. Pearson’s own student Udny Yule used 8 degrees of freedom for a 3×3 table in 1906; Greenwood and Yule noticed something was wrong in 1915 without being able to say what; and it was Fisher, in the 1922 paper that introduced the term “degrees of freedom” into statistics, who explained that estimating the expected frequencies from the marginal totals uses up some of the freedom in the table (Fisher, 1922).
Pearson did not accept the correction. He replied in Biometrika the same year arguing that Fisher had blundered, and the dispute became one of the more bitter in the history of statistics (Stigler, 2008). The consequence for anyone marking a script today is simple: an inflated degrees of freedom means an inflated critical value, which makes it harder to reach significance and can turn a genuine effect into a non-significant result. If you are ever unsure whether you have counted correctly, remember that the totals row and totals column are never counted.
How to read a chi-squared critical values table
To read a chi-squared critical values table, find the row matching your degrees of freedom, then move across to the column for your significance level and the number of tails in your hypothesis. The number where they meet is the critical value your calculated chi-squared must equal or exceed.
Four things must be right before you can pick a number out of the table, and each of them is a place marks are lost.
- The degrees of freedom. Work these out from the shape of the table, not the sample size.
- The significance level. In psychology this is p = 0.05 unless there is a stated reason to be stricter, as explained in our guide to statistical significance.
- One-tailed or two-tailed. This follows from whether your hypothesis was directional.
- The direction of the comparison. Chi-squared must be equal to or greater than the critical value.
One-tailed or two-tailed for chi-squared?
Use the two-tailed column when your hypothesis is non-directional and the one-tailed column when it is directional, exactly as with any other test. At df = 1 and p = 0.05 the two-tailed critical value is 3.84 and the one-tailed value is 2.71.
This is the part of chi-squared that confuses almost everyone, and it is worth being honest about why. The chi-squared distribution genuinely has only one tail. Because every discrepancy is squared, chi-squared can never be negative, and a large value always means the same thing: the data have departed from the null hypothesis. There is no “other end” of the distribution for a result to fall into. So in a strict mathematical sense every chi-squared test is one-tailed.
What the columns in a psychology table are actually doing is something different. Look at the published table and you will see that the one-tailed and two-tailed headings are simply the same six critical values labelled twice, with the one-tailed probability always half the two-tailed one. The value 3.84 sits under “two-tailed 0.05” and also under “one-tailed 0.025”. The table is not offering two different statistical procedures. It is offering the same probability under two labels, so that a student with a directional hypothesis can claim credit for having predicted the direction in advance.
For exam purposes the rule is unambiguous, so use it without agonising: directional hypothesis means the one-tailed column, non-directional means the two-tailed column. Both AQA and Pearson Edexcel print the table with the two rows of headings and expect students to pick the right one (AQA, 2015; Pearson Edexcel, 2026). Just be aware that if you go on to study statistics at university, you will be told that chi-squared has no tails to choose between, and both statements are correct within their own conventions.
Always use the table your board gives you
The published tables do not entirely agree with one another at the strictest levels, so use the one printed in your own exam paper rather than any other, including the one above.
At the 0.05 level every version agrees: 2.71 one-tailed and 3.84 two-tailed at df = 1. Beyond that they diverge. Pearson Edexcel’s Formulae and Statistical Tables booklet for summer 2026 pairs a one-tailed 0.01 with a critical value of 5.02 at df = 1, because that column’s real tail area is 0.025 and the one-tailed heading has been rounded from 0.0125 to 0.01. OCR’s sample H567/01 paper prints a one-tailed 0.01 critical value of 5.41, which is the value for a tail area of exactly 0.02.
Neither is an error. They are two different rounding conventions applied to the same underlying distribution, and each board marks against its own booklet. In practice it almost never matters, because A level questions are set at the 0.05 level where the tables agree. But it is a good reason never to memorise critical values, and never to bring a table from one board into an exam set by another.
How to write up a chi-squared result
A chi-squared result is written in APA format as the symbol, the degrees of freedom and sample size in brackets, the calculated value to two decimal places, and the probability. It looks like this: χ²(1, N = 120) = 8.00, p < .05.
Each element is doing a job. The number in brackets before the comma is the degrees of freedom. N is the total number of participants, not the number per group. The calculated value follows to two decimal places. The probability statement uses a leading decimal point with no zero in front of it, which is the APA convention for any statistic that cannot exceed 1.
In a full report the statistic never stands alone. It comes after a sentence of description and before a sentence of interpretation, so that a reader who skips the numbers still understands the finding.
“Cooperative play was recorded for 45 of the 60 children observed outdoors (75 per cent) compared with 30 of the 60 observed indoors (50 per cent). A chi-squared test found this difference to be significant, χ²(1, N = 120) = 8.00, p < .05. Children in the outdoor setting were therefore substantially more likely to engage in cooperative play.”
Reporting percentages alongside the raw frequencies is good practice and takes one line. The frequencies are what the test used; the percentages are what a reader can interpret. The same principle applies to any results section, as described in our overview of quantitative and qualitative analysis.
Assumptions of the chi-squared test
The chi-squared test rests on four assumptions: the data are frequencies rather than percentages or scores, the observations are independent, the categories are mutually exclusive and exhaustive, and the expected frequencies are large enough for the approximation to hold.
Frequencies, not percentages. Chi-squared must be calculated on raw counts. Running it on percentages is a genuine error rather than a stylistic one, because the test’s sensitivity depends on how many people the counts represent. Forty per cent of 10 people and 40 per cent of 1,000 people carry wildly different amounts of evidence, and a percentage hides that difference completely.
Independence. Each observation must come from a different participant and land in one cell. This is the assumption broken most often, usually by testing the same people twice.
Mutually exclusive and exhaustive categories. Every participant must fit into exactly one category, and there must be a category for everyone. If children could be coded as both cooperative and solitary, or if some children fitted neither, the table would not be valid.
Large enough expected frequencies. This is the assumption with an actual number attached, and it is the subject of the next section.
The expected frequency of 5 rule
The usual rule is that no more than 20 per cent of expected frequencies should fall below 5, and none should fall below 1. In a 2×2 table, where there are only four cells, that means every expected frequency should reach 5.
The rule comes from Cochran (1954), who was trying to say when the chi-squared distribution is a good enough approximation to the true distribution of the test statistic. Chi-squared is a continuous curve; the data are whole numbers of people. When the counts are large the mismatch does not matter. When they are small it does, and the test starts reporting significance more often than it should.
The important word is expected. An observed frequency of zero is perfectly acceptable and does not violate anything. It is the expected frequency, calculated from the margins, that has to clear the threshold. Students regularly abandon a valid analysis because one observed cell was empty, which is not what the rule says.
It is also worth knowing that the rule is conservative. Kroonenberg and Verbeek (2018) reviewed the evidence and concluded that Cochran’s threshold is stricter than it needs to be in many realistic tables, and that analyses are often abandoned unnecessarily. For A level purposes, apply the rule as taught; for a university project, it is worth reading further before discarding data.
What to do when expected frequencies are too small
There are three standard responses when expected frequencies fall below 5, and which one is appropriate depends on why the cells are thin.
- Combine categories. If two categories are conceptually close and both sparse, merging them is usually the cleanest fix. Combining insecure-avoidant and insecure-resistant into a single “insecure” category turns a 2×3 table into a 2×2 one with healthier cells. The decision must be made on theoretical grounds and reported, never chosen after seeing which merge gives a significant result.
- Collect more data. Expected frequencies scale with the sample, so a larger sample solves the problem directly. This is the right answer when the categories all matter and none can sensibly be merged.
- Use a different test. For a 2×2 table with small numbers, Fisher’s exact test computes the probability directly rather than approximating it, and is what statistical software will usually recommend. It is beyond A level but worth knowing the name.
Yates’s correction for continuity
Yates’s correction is an adjustment applied to 2×2 tables that subtracts 0.5 from each absolute difference before squaring, producing a slightly smaller and more conservative chi-squared value. Yates proposed it in 1934 to compensate for using a continuous distribution on whole-number data (Yates, 1934).
Applied to the worked example, each difference of 7.5 becomes 7.0, and the chi-squared value drops from 8.00 to 6.97. In this case the conclusion is unchanged, because 6.97 still exceeds 3.84 comfortably. In a borderline case it could easily flip the decision, which is exactly why statisticians argue about whether it should be used at all: many now regard it as too conservative for tables with adequate expected frequencies.
Neither AQA nor Pearson Edexcel requires Yates’s correction at A level, and you should not apply it in an exam unless a question explicitly asks for it. Know what it is, know it applies only to 2×2 tables, and move on.
Goodness of fit versus test of independence
There are two chi-squared tests, and A level psychology almost always means the second. A goodness of fit test compares one row of observed frequencies against a set of expected proportions you specify in advance. A test of independence compares two variables in a contingency table to see whether the split across one depends on the other.
| Goodness of fit | Test of independence | |
|---|---|---|
| Shape of data | One row of categories | A table of rows and columns |
| Expected values | Come from a theory or a known distribution | Calculated from the row and column totals |
| Degrees of freedom | Categories − 1 | (rows − 1) × (columns − 1) |
| Question answered | Do the data match what was predicted? | Are the two variables related? |
| Typical A level use | Rare | Almost always this one |
A goodness of fit example: if a researcher believed that children’s choice of play area should be evenly spread across four areas, the expected frequency for each would be the total divided by four, and the degrees of freedom would be 4 − 1 = 3. The expectation comes from the researcher’s theory, not from the data.
If an exam question gives you a table with rows and columns, it is a test of independence and the (rows − 1) × (columns − 1) formula applies. If it gives you a single row of counts and tells you what proportions to expect, it is goodness of fit and the degrees of freedom are one less than the number of categories.
Effect size: how big is the association?
A significant chi-squared tells you an effect is unlikely to be chance; it does not tell you how large the effect is. For that you need an effect size, and for contingency tables the standard measure is Cramer’s V, which rescales chi-squared onto a range from 0 to 1.
Cramer’s V is the square root of chi-squared divided by the sample size times the smaller of (rows − 1) and (columns − 1). For the first worked example: 8.00 divided by (120 × 1) is 0.0667, and the square root of that is 0.26. For the second: 18.70 divided by (200 × 1) is 0.0935, giving 0.31.
Cohen’s (1988) conventional benchmarks for this family of effect sizes are 0.1 for a small effect, 0.3 for a medium one and 0.5 for a large one. So the play-setting result is a small to medium effect and the attachment result is a medium one. Both are significant; neither is enormous.
This distinction is the single most useful thing to say in an evaluation question about any inferential test. Significance depends heavily on sample size: run the play study on 1,200 children instead of 120 with the same proportions and chi-squared becomes 80.00, wildly significant, while Cramer’s V stays at exactly 0.26 because the pattern in the data has not changed at all. Significance answers “is this real?” and effect size answers “does it matter?” The two are not the same question, and only one of them is affected by how many people you tested.
Effect size is not required for chi-squared at A level, but knowing that it exists and what it does will earn credit in an evaluation. The wider argument about significance testing, statistical power and error rates is covered in our guide to p-values, errors and effect sizes.
Common chi-squared mistakes in A level exams
Most marks lost on chi-squared questions come from a small number of predictable errors, and all of them are avoidable with a checklist.
- Comparing in the wrong direction. Chi-squared must be equal to or greater than the critical value. Wilcoxon, Mann-Whitney and the sign test all go the other way. If you have memorised one rule for all the tests, you have memorised a mistake.
- Counting the totals as a row or column. A 2×2 table drawn with its margins looks like a 3×3 table. Degrees of freedom come from the data cells only.
- Using percentages instead of frequencies. Chi-squared needs raw counts. Converting to percentages destroys the information about sample size that the test depends on.
- Calculating expected frequencies from the wrong totals. Each cell uses its own row total and its own column total. Using the same pair of totals for every cell is the commonest arithmetic slip, and the row-check in step 2 catches it.
- Rounding too early. Keep expected frequencies to at least two decimal places all the way through. Rounding 37.5 to 38 at the start shifts the final answer enough to matter in a borderline case.
- Choosing chi-squared for repeated measures. Nominal data plus the same participants tested twice is not a chi-squared design, however tempting the frequencies look.
- Stopping at “significant”. A conclusion that does not say what actually happened in the data is an incomplete answer. Name the direction.
- Confusing the symbol. The test statistic is chi-squared, not chi. There is no square root at the end.
Working through past questions with these in mind is far more efficient than re-reading notes, for the reasons set out in our guide to evidence-based revision methods.
What each exam board requires
The boards differ in how much chi-squared arithmetic they expect, and it is worth knowing which applies to you before deciding how much of this article to memorise.
AQA. The specification lists chi-squared among the tests students must know “when to use”, alongside Spearman’s rho, Pearson’s r, Wilcoxon, Mann-Whitney and the two t-tests. Calculation is specified only for the sign test. What AQA does require for every test is “use of statistical tables and critical values in interpretation of significance”, so you must be able to read the table, work out the degrees of freedom and reach a conclusion (AQA, 2015).
Pearson Edexcel. More is required here. The specification asks for “reasons for choosing a chi-squared test; comparing observed and critical values to judge significance; the chi-squared test”, and the practical research exercise expects students to “analyse the findings to produce results, including using a chi-squared test”. The formula and the full critical values table are printed in the specification’s appendix and provided in the exam (Pearson Edexcel, 2026).
OCR. Chi-squared is examined in the H567/01 Research Methods paper. The sample assessment material asks for the reasons behind choosing the test, how the degrees of freedom are determined, and a judgement of significance against critical values printed in the question itself, rather than for the full calculation (OCR, 2026).
The honest advice is to learn the calculation whichever board you sit. It takes twenty minutes to learn and it makes the “when to use it” questions far easier to answer, because you can see what the test is doing rather than recalling a rule about it.
How chi-squared compares with the other tests
Chi-squared is one of seven or eight tests on the A level specifications, and the choice between them comes down to three questions: what are you testing for, what level of measurement are your data, and what was the experimental design?
| Test | Looking for | Level of measurement | Design |
|---|---|---|---|
| Chi-squared | Difference or association | Nominal | Independent groups |
| Sign test | Difference | Nominal | Repeated measures |
| Mann-Whitney U | Difference | Ordinal | Independent groups |
| Wilcoxon | Difference | Ordinal | Repeated measures |
| Unrelated t-test | Difference | Interval | Independent groups |
| Related t-test | Difference | Interval | Repeated measures |
| Spearman’s rho | Correlation | Ordinal | Pairs of scores |
| Pearson’s r | Correlation | Interval | Pairs of scores |
Chi-squared is also the only test in this set that handles more than two categories in a single analysis without any extra machinery. A 2×4 table is no harder than a 2×2 one; only the degrees of freedom change. Comparing four groups with t-tests would mean six separate comparisons and a multiple comparisons problem.
The full reasoning for picking between all eight, with the decision tree laid out, is in our guide to choosing the right statistical test. For the parametric alternatives specifically, see our guides to t-tests and the standard deviation that underpins them.
Where chi-squared appears in real psychology
Chi-squared turns up wherever psychologists code behaviour into categories rather than measuring it on a scale, which is a much larger slice of the discipline than the statistics chapter suggests.
Observational research is the obvious home. Whenever a behavioural coding scheme produces tallies in named categories, and those tallies are compared across two or more groups, chi-squared is the test. The same is true of content analysis, where the units counted are references in a text rather than behaviours in a room.
Attachment research is a good example of the shape of data chi-squared handles. The Strange Situation produces a classification, not a score, and van IJzendoorn and Kroonenberg’s (1988) meta-analysis of 32 samples across 8 countries was essentially an exercise in comparing distributions of those classifications. Their central finding was that variation within countries was larger than variation between them, which is a conclusion about how the frequencies are distributed rather than about any average. Our article on Schaffer and Emerson’s stages of attachment describes another study built on categorical observations.
Social influence research uses it too. Studies that record whether each participant conformed or did not, obeyed or refused, produce exactly the two-category outcome that chi-squared was built for, which is why it appears throughout the literature covered in our overview of social influence and in work on resistance to social influence. Case studies, by contrast, rarely use it at all, for reasons our guide to the case study method explains: a single participant gives you no frequencies to compare.
Conclusion
The chi-squared test asks one question: are the frequencies you observed far enough from the frequencies you would expect by chance to be worth taking seriously? Everything else is bookkeeping. Expected frequencies come from the row and column totals, the differences are squared so they accumulate rather than cancel, dividing by the expected frequency keeps every cell in proportion, and the total is compared against a table.
Three details account for most of the marks lost. Degrees of freedom come from the shape of the table, never the sample size, and never include the totals. The comparison runs the opposite way from the sign test, Wilcoxon and Mann-Whitney: chi-squared must be equal to or greater than the critical value. And a conclusion is not finished until it says what actually happened in the data, not merely that something did.
If the test itself now makes sense, the next useful step is deciding when to reach for it rather than something else, which is the job of the decision guide linked throughout this article. And if you want to see the method working, take the OCR data above, put it into the calculator, and watch 3.80 come out.
Frequently Asked Questions
What is the chi-squared test used for?
The chi-squared test is used to find out whether two categorical variables are related. It compares the number of people who actually fell into each category with the number you would expect if the variables were unconnected. Psychologists use it for observational coding, content analysis, and any study where the outcome is which box a participant went into rather than how much of something they produced.
How do you calculate degrees of freedom for chi-squared?
Multiply the number of data rows minus one by the number of data columns minus one. A 2×2 table gives 1 degree of freedom, a 2×3 table gives 2, and a 3×4 table gives 6. Never include the row of totals or the column of totals in the count, and never use the number of participants: degrees of freedom depend only on the shape of the table.
Is the chi-squared test parametric or non-parametric?
Chi-squared is non-parametric. It makes no assumption that the underlying population is normally distributed, and it works on counts rather than on measured scores, so there is no mean or standard deviation for it to depend on. That is what allows it to be used on nominal data, the weakest level of measurement, where parametric tests cannot go.
What is the critical value of chi-squared at 0.05?
It depends on the degrees of freedom and on whether the hypothesis is directional. For a two-tailed test at the 0.05 level the critical values are 3.84 at df 1, 5.99 at df 2, 7.82 at df 3 and 9.49 at df 4. For a one-tailed test at the same level they are 2.71, 4.61, 6.25 and 7.78. The full table is earlier in this article.
What is the difference between a chi-squared test and a t-test?
They work on different kinds of data. A t-test compares the means of two sets of measured scores and needs interval data that are roughly normally distributed. Chi-squared compares counts of people in named categories and needs no scores at all. If your results are averages, use a t-test. If they are tallies, use chi-squared.
Why must expected frequencies be more than 5?
Because the chi-squared distribution is a smooth curve being used to approximate the behaviour of whole numbers of people. When cells are thin the approximation drifts and the test reports significance too readily. Cochran’s (1954) rule allows up to a fifth of cells below 5 and none below 1. The rule applies to expected frequencies only, never to observed ones, so an empty observed cell is not a problem.
Can you use chi-squared for repeated measures?
No. Chi-squared requires that every observation is independent, meaning each participant appears in exactly one cell. Testing the same people twice breaks that requirement, however the data are arranged. For nominal data from a repeated measures design, the sign test is the standard choice at A level, and McNemar’s test is the equivalent used beyond it.
How is chi-squared pronounced?
Chi rhymes with “sky” and “pie”: say “kye-squared”. The word is Greek, and the letter is written as a lower-case x. Saying it as “chee” or “chai” is a common mistake in classrooms and it will not lose you any marks, but you will hear it said the other way in every university lecture.
What does a significant chi-squared result actually mean?
It means the pattern of frequencies in your table is unlikely to have arisen by chance if the two variables were unrelated. It does not say how strong the relationship is, which needs an effect size such as Cramer’s V, and it does not say which way round the relationship runs. For direction, look back at the original frequencies and describe what you see.
References
AQA. (2015). AS and A-level Psychology specification (7181, 7182). Assessment and Qualifications Alliance.
Cochran, W. G. (1954). Some methods for strengthening the common chi-squared tests. Biometrics, 10(4), 417-451.
Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates.
Cramer, H. (1946). Mathematical methods of statistics. Princeton University Press.
Fisher, R. A. (1922). On the interpretation of chi-squared from contingency tables, and the calculation of P. Journal of the Royal Statistical Society, 85(1), 87-94.
Greenwood, M., & Yule, G. U. (1915). The statistics of anti-typhoid and anti-cholera inoculations, and the interpretation of such statistics in general. Proceedings of the Royal Society of Medicine, Section of Epidemiology and State Medicine, 8, 113-190.
Kroonenberg, P. M., & Verbeek, A. (2018). The tale of Cochran’s rule: My contingency table has so many expected values smaller than 5, what am I to do? The American Statistician, 72(2), 175-183.
OCR. (2026). A Level Psychology H567/01 Research methods: Sample question paper and mark scheme. Oxford Cambridge and RSA Examinations.
Pearson Edexcel. (2026). Specification: Pearson Edexcel Level 3 Advanced GCE in Psychology (9PS0) (Issue 4). Pearson Education Limited.
Pearson, K. (1900). On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling. The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science, 50(302), 157-175.
Stigler, S. M. (2008). Karl Pearson’s theoretical errors and the advances they inspired. Statistical Science, 23(2), 261-271.
van IJzendoorn, M. H., & Kroonenberg, P. M. (1988). Cross-cultural patterns of attachment: A meta-analysis of the strange situation. Child Development, 59(1), 147-156.
Yates, F. (1934). Contingency tables involving small numbers and the chi-squared test. Supplement to the Journal of the Royal Statistical Society, 1(2), 217-235.
Yule, G. U. (1906). The influence of bias and of personal equation in statistics of ill-defined qualities. Journal of the Anthropological Institute of Great Britain and Ireland, 36, 325-381.
Further Reading and Research
Recommended Articles
- Choosing the Right Statistical Test: A Decision Guide
- Levels of Measurement: Nominal, Ordinal, Interval
- The Sign Test: Formula, Worked Example and Critical Values
Suggested Books
- Discovering Statistics Using IBM SPSS Statistics by Andy Field
- The standard undergraduate text, and the friendliest serious statistics book in print. The chapter on categorical data covers chi-squared, Fisher’s exact test, the loglinear models that generalise it, and how to report all of them.
- Statistics Without Tears: An Introduction for Non-Mathematicians by Derek Rowntree
- Almost no arithmetic and no formulae at all. Useful if the concepts behind significance and sampling have never quite settled, because it explains the reasoning before anything is calculated.
- The Lady Tasting Tea: How Statistics Revolutionised Science in the Twentieth Century by David Salsburg
- The story of the people rather than the methods, including Pearson, Fisher and the argument between them that produced the modern understanding of degrees of freedom.
Recommended Websites
- AQA A-level Psychology specification
- The primary source for what is actually examinable. The research methods section sets out exactly which tests must be known and what has to be done with them.
- Pearson Edexcel Psychology qualification pages
- The Edexcel specification prints the chi-squared formula and the complete critical values table in its appendix, which is the same table reproduced in this article.
- Project Euclid
- Hosts Statistical Science, including Stigler’s open-access account of Pearson’s degrees of freedom error and Fisher’s correction.
To cite this article please use:
Early Years TV Chi-Squared Test in Psychology: Formula, Example, Table. Available at: https://www.earlyyears.tv/chi-squared-test-psychology-formula-example-table/ (Accessed: 2 October 2026).

