Which Statistical Test to Use in Psychology: A Decision Guide

In a review of 513 neuroscience papers, of the studies where a particular comparison mattered, 79 used the wrong statistical procedure and only 78 used the right one (Nieuwenhuis et al., 2011). Knowing which statistical test to use is not a beginner’s problem.
Key Takeaways
- Three questions decide the test: is the hypothesis predicting a difference or a relationship, is the design related or unrelated, and is the data nominal, ordinal or interval.
- Those three answers point to exactly one cell of a three-by-three table containing all eight tests named in the AQA specification, so the choice is mechanical once you have them.
- Most marks are lost on the second and third questions, not the first: matched pairs is a related design, and a rating scale is ordinal unless the units are real and equal.
Choosing a statistical test feels like the hardest part of research methods, and it is the part students most often get wrong under exam conditions. The information needed is never hidden. It is always sitting in the question stem, in the hypothesis and the description of the method. The difficulty is that the test names arrive as a list of eight unfamiliar words, and a list is very hard to search through when you are tired and the clock is running.
The list is the wrong shape. The eight tests are not a list at all, they are a grid, and the grid has only two dimensions. Once you can see the grid, the choice stops being an act of recall and becomes an act of reading. You read the hypothesis, you read the method, you read off the cell. This guide builds that grid from the three questions the specification actually names, then tests it against six scenarios of the kind that appear on exam papers.
You will get more out of this if you are already comfortable with what a significant result means, since choosing the test is only the first half of the job. Our guide to statistical significance, p-values and errors covers the second half. If you have met only one test so far, it was almost certainly the sign test, and that is a good place to anchor everything that follows.
How Do You Know Which Statistical Test to Use in Psychology?
You know which statistical test to use in psychology by answering three questions about the study, in a fixed order. Is the hypothesis predicting a difference between conditions, or a relationship between two co-variables? Is the design related or unrelated? And at what level is the data measured: nominal, ordinal or interval? Three answers identify one test, every time.
This is not a revision shortcut invented by teachers. It is the specification. The AQA A-level Psychology specification lists, under inferential testing, “factors affecting the choice of statistical test, including level of measurement and experimental design”, and then names the tests students must be able to place: Spearman’s rho, Pearson’s r, Wilcoxon, Mann-Whitney, related t-test, unrelated t-test and Chi-Squared, alongside the sign test introduced earlier (AQA, 2021). Eight tests, three questions, one answer.
The order matters more than students expect. The first question splits the eight tests into two families that behave completely differently, and getting it wrong sends you down a branch from which no amount of care about the other two questions can recover you. The second question only applies to one of those families. The third question is the one that actually decides between similar-looking tests, and it is where most of the marks are quietly lost.
It is worth being honest about why this matters beyond the exam. Statistical reporting in published psychology is not in good shape. Analysing more than 250,000 p-values from eight major psychology journals across three decades, Nuijten et al. (2016) found that half of all published psychology papers using significance testing contained at least one p-value inconsistent with its own test statistic and degrees of freedom, and that one paper in eight contained an inconsistency large enough to have changed the conclusion. Learning to choose a test carefully is a habit that keeps paying.

The Statistical Test Table Every Psychology Student Needs
The statistical test table arranges all eight tests into three rows and three columns. The rows are the three levels of measurement, from least informative at the top to most informative at the bottom. The columns are what the hypothesis is looking for: a difference in an unrelated design, a difference in a related design, or a relationship between two co-variables.
| Level of measurement | Difference, unrelated design | Difference, related design | Relationship between co-variables |
|---|---|---|---|
| Nominal | Chi-Squared | Sign test | Chi-Squared |
| Ordinal | Mann-Whitney U | Wilcoxon signed-ranks | Spearman’s rho |
| Interval | Unrelated t-test | Related t-test | Pearson’s r |
Reading the table in the right order
Read the table across before you read it down. Choosing a column means settling what kind of claim the hypothesis is making and how participants were allocated, which are facts about the design and are stated plainly in any exam question. Choosing a row means judging the data, which involves more interpretation. Settling the easier question first leaves you with a shortlist of three rather than eight.
The rows also carry a hidden message about power. Moving down the table, each row uses more of the information in the data. Nominal tests know only which category a score fell into. Ordinal tests know the order of the scores. Interval tests know the actual distances between them. A test that uses more information is more likely to detect a real effect, which is why you should never deliberately throw information away by treating interval data as though it were ordinal.
Why Chi-Squared appears twice
Chi-Squared appears twice because it is the only test in the table that works as both a test of difference and a test of association. With nominal data, asking whether two groups differ in how they are distributed across categories and asking whether two categorical variables are associated turn out to be the same arithmetic question, so the same test answers both (Pearson, 1900).
This is genuinely useful rather than a quirk to memorise. If you are working with counts of people in categories and you cannot decide whether the hypothesis is about difference or association, Chi-Squared is correct either way. The nominal row is the only row where the first question can be answered wrongly without costing you the test, and it is worth knowing that safety net is there.
The one exception in that row is the sign test, which sits in the related-design cell. It is the only nominal test that requires the same participants to be measured twice, because it works on whether each person’s score went up or down (Coolican, 2018).
Statistical Test Picker
Answer the three questions below and the picker returns the test, a note on what the test actually does, and a sentence in the wording examiners expect for a justification. The full table sits underneath it, so you can check the answer against the grid and start learning the pattern rather than relying on the tool.
Which Statistical Test Should I Use?
Answer the three questions the AQA specification names. The picker gives you the test and a sentence you can use to justify it.
| Data level | Unrelated difference | Related difference | Relationship |
|---|---|---|---|
| Nominal | Chi-Squared | Sign test | Chi-Squared |
| Ordinal | Mann-Whitney U | Wilcoxon signed-ranks | Spearman’s rho |
| Interval | Unrelated t-test | Related t-test | Pearson’s r |
Question One: Is the Hypothesis Predicting a Difference or a Relationship?
The hypothesis is predicting a difference if it compares two sets of scores against each other, and a relationship if it asks whether two measurements taken from the same people rise and fall together. The clue is in how many things were measured per participant, and whether anything was manipulated.
What a test of difference looks like
A test of difference compares one condition with another. There is an independent variable that the researcher has set up or selected, and a dependent variable that gets measured in each condition. The hypothesis contains a comparative word: higher, lower, faster, more, fewer, better.
For example: “Participants who revise in silence will recall more words than participants who revise with music playing.” Two conditions, one measurement each, a comparison. That is a test of difference, and it sits in one of the first two columns.
What a test of relationship looks like
A test of relationship takes two measurements from every participant and asks how they move together. There is no independent variable and no conditions, only two co-variables. Nothing is manipulated, which is why a correlational study can never on its own establish that one thing causes another.
For example: “There will be a relationship between the number of hours a student revises and the mark they achieve.” Every student gives you an hours figure and a marks figure. That is a test of relationship, and it sits in the third column. If you want the detail on what the resulting number means, our guide to correlation coefficients and Pearson’s r works through the interpretation.
The trap: a correlation dressed up as a difference
The commonest error on this question is treating two measurements from the same person as two conditions. A study measuring anxiety and sleep quality in fifty adults has two scores per person, which looks superficially like a repeated measures design. It is not. Nobody was put into a condition and nothing was manipulated, so it is a correlation.
The test is simple. Ask whether the two numbers are measuring the same thing on the same scale. In a repeated measures experiment they are: recall score before and recall score after, both out of twenty. In a correlation they are not: hours of sleep and an anxiety rating are different quantities that could not sensibly be subtracted from one another. If subtracting one from the other would be meaningless, you are looking at a correlation.
Question Two: What Is the Experimental Design?
The experimental design tells you whether the two sets of scores are paired up or not. If each score in one condition is tied to a specific score in the other, the design is related. If the two conditions contain different, unconnected people, the design is unrelated. This question only applies to tests of difference, because a correlation is paired by definition.
Independent groups is an unrelated design
Independent groups means every participant experiences one condition only, so the two sets of scores come from different people. There is no way to say which score in condition A belongs with which score in condition B, and the two groups need not even be the same size.
That lack of pairing is what forces an unrelated test. Mann-Whitney U and the unrelated t-test both work by comparing two whole groups rather than by looking at individual pairs (Mann and Whitney, 1947).
Repeated measures is a related design
Repeated measures means every participant takes part in both conditions, so each person gives you a pair of scores. The pairing is what a related test exploits: instead of comparing two groups, it works on the difference within each pair, which strips out the variation caused by people simply being different from one another.
This is why related designs are more sensitive. Wilcoxon’s signed-ranks test ranks those within-person differences by size, and the related t-test averages them (Wilcoxon, 1945; Student, 1908). Both are looking at change within a person rather than at a gap between two crowds.
Matched pairs counts as related, not unrelated
Matched pairs is a related design, even though every participant appears in only one condition. This is the single most expensive misreading in the whole topic, and it is easy to see why it happens: the words “different participants in each condition” are true of matched pairs and are also the definition of independent groups.
What makes it related is that the researcher has deliberately paired each participant with a specific other participant on the variables that matter, such as age, IQ or baseline anxiety. Those two people are then split across the conditions. Score number seven in condition A has a designated partner in condition B, so the scores come in pairs and a related test applies.
A one-line rule covers all three designs. Can you draw a line from each score in one condition to exactly one score in the other, and defend why those two belong together? If yes, the design is related, whether the pairing comes from the same person or a matched partner. If no, it is unrelated.
Question Three: What Is the Level of Measurement?
The level of measurement describes how much information a number carries. Nominal data only says which category something fell into. Ordinal data puts things in order but the gaps between positions are not equal. Interval data has real, equal units, so the gap between 10 and 20 means the same as the gap between 40 and 50. The AQA specification names these three and stops there (AQA, 2021).
This scheme comes from Stevens (1946), who set out four scales of measurement rather than three. The fourth, ratio, is interval data with a true zero, so weight and reaction time are ratio while temperature in Celsius is only interval. For test selection the distinction makes no difference, which is why A-level treats interval and ratio together and why the picker above does the same.
Nominal data
Nominal data is data in named categories, where the only number involved is a count of how many fell into each one. Eye colour, whether a child was securely or insecurely attached, whether someone passed or failed, which of three posters a shopper chose: all nominal. There is no sense in which one category is higher than another.
The recognition test is whether you are counting people or measuring them. “Thirty-one participants conformed and nineteen did not” is nominal, because 31 and 19 are headcounts. “Participants conformed on an average of 4.2 trials out of 12” is not nominal, because 4.2 is a measurement of a person, not a count of people.
Ordinal data
Ordinal data can be put in a meaningful order, but the distances between the positions are unknown or uneven. Finishing positions in a race are the clearest case: first, second and third tell you the order but nothing about the gaps, which might be a tenth of a second and then four minutes.
In psychology, most rating scales fall here. A participant rating their anxiety from 1 to 10 is producing ordinal data, because there is no guarantee that the step from 3 to 4 is the same size as the step from 8 to 9. The same applies to Likert-scale agreement items, to a teacher’s ranking of children by confidence, and to any score derived from a questionnaire with no established units.
Interval data
Interval data is measured in units that are real and equal all the way along the scale. Time in seconds, temperature in degrees, number of words correctly recalled, heart rate in beats per minute, score on a standardised IQ test: each unit means the same thing wherever it sits on the scale.
Interval data unlocks the bottom row, which contains the parametric tests. These use the mean and the spread of the scores rather than just their order, which is why they can detect smaller effects from the same number of participants. Our guides to t-tests and to standard deviation cover how that extra information gets used.
Deciding between ordinal and interval under exam pressure
Ask one question: does one unit on this scale mean the same thing everywhere on it? If yes, treat it as interval. If you cannot defend that, treat it as ordinal. Almost every disputed case in an exam paper is a rating scale or a questionnaire total, and both default to ordinal.
| Measurement | Level | Why |
|---|---|---|
| Number of words recalled from a list of 20 | Interval | Each word is one unit and every unit is identical |
| Self-rated stress from 1 to 10 | Ordinal | Ordered, but the steps are not guaranteed equal |
| Number of children who chose the aggressive toy | Nominal | A count of people in a category |
| Reaction time in milliseconds | Interval | Real units with a true zero, so also ratio |
| Rank order of preference for five adverts | Ordinal | Positions only, with unknown gaps |
| Attachment type from the Strange Situation | Nominal | Named categories with no order |
There is a live argument among researchers about whether the ordinal default is too cautious. Norman (2010) reviewed evidence going back to the 1930s and concluded that parametric methods are robust to violations of their assumptions, including when applied to Likert data, so the standard objection that ordinal data rules out parametric tests is overstated. That debate is worth knowing about, but it does not change what an A-level mark scheme expects: rating scales are ordinal.
The Carrots Mnemonic for Remembering the Statistical Test Table
The best-known mnemonic for the statistical test table is “Carrots Should Come Mashed With Swede Under Roast Potatoes”. Each initial letter is a test, and the nine tests are read across the table row by row, starting at the top left.
| Word | Test | Position in the table |
|---|---|---|
| Carrots | Chi-Squared | Nominal, unrelated difference |
| Should | Sign test | Nominal, related difference |
| Come | Chi-Squared | Nominal, relationship |
| Mashed | Mann-Whitney U | Ordinal, unrelated difference |
| With | Wilcoxon | Ordinal, related difference |
| Swede | Spearman’s rho | Ordinal, relationship |
| Under | Unrelated t-test | Interval, unrelated difference |
| Roast | Related t-test | Interval, related difference |
| Potatoes | Pearson’s r | Interval, relationship |
The mnemonic only works if you also remember the frame it fills. Nine words in a row are useless without knowing that the rows run nominal, ordinal, interval and the columns run unrelated, related, relationship. Write the empty grid first, then fill it from the phrase. Two minutes at the start of an exam buys you the whole table on the page in front of you, which is worth far more than trying to recall nine names cold under pressure.
It also has a pleasing internal logic that helps it stick. Every row keeps its initial letter: the ordinal row is M, W, S for Mann-Whitney, Wilcoxon and Spearman; the interval row is U, R, P for unrelated, related and Pearson. Only the nominal row breaks the pattern, and it breaks it by repeating Chi-Squared, which is the one irregularity in the table anyway.
Worked Examples: Choosing a Test in Six Scenarios
The three questions are quick once you have practised them. Each scenario below is answered in the same order every time: difference or relationship, then design, then level of measurement.
Scenario one: noise and recall
Forty participants are randomly allocated to revise either in silence or with background noise, and each recalls as many words as they can from a list of twenty. The answer is the unrelated t-test. It predicts a difference, participants did one condition each so the design is independent groups, and words recalled is interval because every word is one identical unit.
Scenario two: a mindfulness intervention
Twenty adults rate their stress from 1 to 10 before an eight-week mindfulness course and again afterwards. The answer is the Wilcoxon signed-ranks test. It predicts a difference, the same people are measured twice so the design is repeated measures, and a self-rating from 1 to 10 is ordinal.
Scenario three: screen time and sleep
A researcher records daily screen time in minutes and average nightly sleep in minutes for sixty teenagers. The answer is Pearson’s r. Two measurements come from every person with nothing manipulated, so it is a relationship rather than a difference, and minutes are interval. Because design does not apply to a correlation, the second question is skipped entirely.
Scenario four: attachment type and school readiness
A study classifies 120 four-year-olds as securely or insecurely attached, and separately as school-ready or not school-ready, then counts how many fall into each combination. The answer is Chi-Squared. Both variables are named categories and the data are counts of children, which is nominal. Whether you frame it as a difference between attachment groups or an association between two categorical variables, the nominal row gives the same test.
Scenario five: a matched pairs memory study
Thirty pairs of participants are matched on age and IQ, one of each pair is taught a mnemonic strategy and the other is not, and all sixty then rank a set of ten images by how memorable they found them. The answer is the Wilcoxon signed-ranks test. Matched pairs is a related design despite each person appearing once, and ranking produces ordinal data. Answering “Mann-Whitney U” here is the single most common error in the topic.
Scenario six: practitioner confidence ratings
Fifty early years practitioners rate their confidence in supporting speech and language on a five-point scale, and the researcher also records how many years each has worked in the sector. The answer is Spearman’s rho. Two measures per person with nothing manipulated makes it a relationship, and a five-point confidence rating is ordinal, which pulls the whole analysis into the ordinal row even though years of experience on its own would be interval (Spearman, 1904).
That last point generalises usefully. When the two co-variables sit at different levels, the lower level decides the test. Mixing an interval measure with an ordinal one gives you an ordinal analysis, because the weaker measurement limits what the pair of numbers can support.
Parametric and Non-Parametric Tests: What the Distinction Adds
Parametric tests are the three in the bottom row: the unrelated t-test, the related t-test and Pearson’s r. They work with the actual values of the scores and make assumptions about the population those scores came from. The other six are non-parametric: they work with counts or ranks and make far fewer assumptions (Siegel, 1956).
The distinction is not a fourth question. If you have already established that the data are interval, you are in the parametric row, and if you have not, you are not. What the distinction adds is an explanation of why the table is shaped the way it is, and a vocabulary for evaluation questions.
The three parametric assumptions
Parametric tests make three assumptions. The data must be at the interval level. The population the sample came from should be roughly normally distributed. And the two conditions should have similar variances, which is called homogeneity of variance. If all three hold, a parametric test is the better choice because it uses more of the information available.
The pay-off is statistical power, meaning the probability of detecting an effect that is genuinely there. A more powerful test needs fewer participants to reach the same conclusion, which matters when recruiting is difficult (Cohen, 1992). The relationship between sample size, variability and how confident you can be in a mean is worked through in our guide to standard error.
When a non-parametric test is the better choice
A non-parametric test is the better choice when the assumptions fail, and not merely an acceptable fallback. Blair and Higgins (1980) compared the Wilcoxon rank-sum statistic with Student’s t across a range of non-normal distributions and found the rank-based test could be substantially more powerful, sometimes by a wide margin. Ranking is not always the weaker option.
Ranks are also far less disturbed by extreme scores. One participant who took forty seconds when everyone else took four will drag a mean a long way, but in a ranking they are simply last. For small, skewed or messy datasets, which describes a great deal of real psychological research, that robustness is worth more than the theoretical power advantage of a t-test.
How to Write the Justification in an Exam Answer
Exam questions rarely ask only for the test name. They ask you to identify the test and justify the choice, usually for three marks, and the mark scheme awards one mark for the test and one for each of the two or three reasons. The efficient answer names the test and then states the three answers as plain facts about the study.
A worked template: “A Mann-Whitney U test should be used. The hypothesis predicts a difference between the two conditions, the study uses an independent groups design because each participant took part in only one condition, and the data are ordinal because participants gave a rating rather than a measurement in fixed units.”
Three habits protect the marks. Name the design in the specification’s own words rather than describing it loosely. Say which level the data are at and add the reason, because “the data are ordinal” alone often earns nothing. And when the question gives you a scenario, quote the detail from that scenario rather than writing a general definition, since the marks are for applying the rule to this study and not for reciting it.
This is worth taking seriously simply because of how the marks are distributed. At least 10% of the marks across A-level Psychology assessments require mathematical skills, and research methods content is assessed on all three papers (AQA, 2021).
Six Common Mistakes When Choosing a Statistical Test
Most wrong answers come from a small set of recurring errors, and all six below are fixable by a habit rather than by more memorising.
- Treating matched pairs as unrelated. Different people in each condition does not mean unpaired. If the researcher deliberately matched them, the design is related.
- Reading two measures per person as two conditions. Anxiety and sleep quality from the same participant is a correlation, not repeated measures. Ask whether subtracting one from the other would mean anything.
- Calling a rating scale interval. A 1 to 10 confidence rating is ordinal. Unless the units are real and equal, stay in the ordinal row.
- Calling a mean score nominal. Nominal data means counting people, not measuring them. An average of 4.2 trials is a measurement.
- Asking about design in a correlation. Related and unrelated do not apply to a test of relationship. Skip straight from question one to question three.
- Choosing the test before reading the hypothesis. The hypothesis states whether the prediction is about difference or relationship. Reading the method first invites you to guess.
A seventh error is subtler and belongs to coursework and undergraduate projects more than exams: choosing the test after seeing the results. Running several tests and reporting the one that reached significance inflates the chance of a false positive, and it is the practice that sits behind a good deal of the replication difficulty psychology has faced. Decide the test from the design, before the data arrive.
What Happens After You Have Chosen the Test
Once the test is chosen, four steps follow. Calculate the test statistic. Work out the value of N or the degrees of freedom. Decide whether the hypothesis was directional or non-directional, which sets whether you read the one-tailed or two-tailed column. Then compare the calculated value with the critical value from the table for that test at p < 0.05.
One detail catches people out. For most of the non-parametric tests, including the sign test, Mann-Whitney and Wilcoxon, the calculated value must be equal to or smaller than the critical value for the result to be significant. For Chi-Squared, the t-tests and the correlation coefficients it is the other way round: the calculated value must be equal to or larger. Every critical value table states which rule applies, and it is always worth reading that line rather than assuming.
The choice of test also shapes how the results get written up. A report states the test, the calculated value, N or the degrees of freedom, the significance level and whether the null hypothesis was rejected. Choosing the test correctly is what makes the rest of that sentence meaningful, and it is the step everything downstream depends on. Where the data are not numerical at all, a different route applies, which our guide to quantitative and qualitative analysis sets out, and single-participant work has its own conventions covered in our guide to case study research methods.
Conclusion
Choosing a statistical test is a three-question procedure, not a feat of memory. Establish whether the hypothesis predicts a difference or a relationship, establish whether the design pairs the scores or not, establish the level of measurement, and the three-by-three table returns one answer. The mnemonic is useful for rebuilding the table under pressure, but it is the frame that does the work, not the phrase.
The errors that cost marks are consistent and few: matched pairs read as unrelated, a correlation read as repeated measures, and a rating scale read as interval. Practising the three questions on scenarios until the order becomes automatic fixes all three, and carries over well beyond the exam into any research where the design has to be settled before the data arrive. Anyone designing a study with participants should read that alongside our guide to psychology research ethics.
Frequently Asked Questions
How do you know which statistical test to use in psychology?
Work through three questions in order. First, does the hypothesis predict a difference or a relationship? Second, if it is a difference, is the design related or unrelated? Third, are the data nominal, ordinal or interval? Those answers identify one cell in the standard three-by-three table of tests, and that cell is the test.
What are the levels of measurement in psychology?
A-level psychology uses three. Nominal means counts of people in named categories, such as how many passed and how many failed. Ordinal means scores that can be put in order but with uneven gaps, such as a rating from 1 to 10. Interval means real, equal units, such as seconds or words recalled. Stevens originally described a fourth, ratio, which is interval with a true zero.
What is the difference between parametric and non-parametric tests?
Parametric tests use the actual values of the scores and assume the data are interval, roughly normally distributed and similar in spread across conditions. Non-parametric tests use counts or ranks instead and assume much less. Parametric tests are more powerful when their assumptions hold; non-parametric tests cope far better when the data are skewed or contain extreme scores.
Is the chi-squared test parametric or non-parametric?
Chi-squared is non-parametric. It works on frequency counts in categories rather than on measured scores, so it makes no assumption that the underlying population is normally distributed. Only three tests at A-level are parametric: the related t-test, the unrelated t-test and Pearson’s r, all of which require interval data.
Should I use ANOVA or chi-squared?
They answer different questions. Chi-squared handles frequency counts in categories. ANOVA compares means across three or more conditions and needs interval data. If your data are headcounts, chi-squared applies. If they are measurements and you have more than two conditions, ANOVA applies, though ANOVA sits outside the A-level specification and appears at undergraduate level.
What is the mnemonic for statistical tests in psychology?
“Carrots Should Come Mashed With Swede Under Roast Potatoes.” The initials give Chi-Squared, Sign test, Chi-Squared, Mann-Whitney, Wilcoxon, Spearman’s rho, Unrelated t-test, Related t-test and Pearson’s r, read across the table row by row from the top left. Sketch the empty grid first, or the nine words have nowhere to go.
When do you use the Mann-Whitney U test?
Use Mann-Whitney U when all three conditions hold: the hypothesis predicts a difference, the design is independent groups, and the data are ordinal. A study comparing how two separate groups of people rated something on a scale is the classic case. If the participants were matched or measured twice, Wilcoxon applies instead.
What statistical tests does A-level psychology cover?
Eight. The specification introduces the sign test at AS and then names seven more for A-level: Spearman’s rho, Pearson’s r, Wilcoxon, Mann-Whitney, the related t-test, the unrelated t-test and chi-squared. Students must be able to say when each is used; full calculation is only required for the sign test.
References
AQA. (2021). AS and A-level Psychology specification (7181, 7182), version 1.2. Assessment and Qualifications Alliance.
Blair, R. C., & Higgins, J. J. (1980). A comparison of the power of Wilcoxon’s rank-sum statistic to that of Student’s t statistic under various nonnormal distributions. Journal of Educational Statistics, 5(4), 309-335.
Cohen, J. (1992). A power primer. Psychological Bulletin, 112(1), 155-159.
Coolican, H. (2018). Research methods and statistics in psychology (7th ed.). Routledge.
Mann, H. B., & Whitney, D. R. (1947). On a test of whether one of two random variables is stochastically larger than the other. The Annals of Mathematical Statistics, 18(1), 50-60.
Nieuwenhuis, S., Forstmann, B. U., & Wagenmakers, E.-J. (2011). Erroneous analyses of interactions in neuroscience: A problem of significance. Nature Neuroscience, 14(9), 1105-1107.
Norman, G. (2010). Likert scales, levels of measurement and the “laws” of statistics. Advances in Health Sciences Education, 15(5), 625-632.
Nuijten, M. B., Hartgerink, C. H. J., van Assen, M. A. L. M., Epskamp, S., & Wicherts, J. M. (2016). The prevalence of statistical reporting errors in psychology (1985-2013). Behavior Research Methods, 48(4), 1205-1226.
Pearson, K. (1900). On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling. Philosophical Magazine, 50(302), 157-175.
Siegel, S. (1956). Nonparametric statistics for the behavioral sciences. McGraw-Hill.
Spearman, C. (1904). The proof and measurement of association between two things. American Journal of Psychology, 15(1), 72-101.
Stevens, S. S. (1946). On the theory of scales of measurement. Science, 103(2684), 677-680.
Student. (1908). The probable error of a mean. Biometrika, 6(1), 1-25.
Wilcoxon, F. (1945). Individual comparisons by ranking methods. Biometrics Bulletin, 1(6), 80-83.
Further Reading and Research
Recommended Articles
- The Sign Test: Formula, Worked Example and Critical Values – the one test the specification asks you to calculate in full, with a calculator and critical value table.
- Statistical Significance: P-values, Errors and Effect Sizes – what happens after the test statistic is calculated, including Type I and Type II errors.
- T-Tests in Psychology: A Complete Guide – the two parametric tests of difference, their assumptions and how they are reported.
Suggested Books
- Research Methods and Statistics in Psychology (7th ed.) by Hugh Coolican
- The standard undergraduate reference, and unusually good on the reasoning behind test selection rather than just the procedures. Works well for A-level students who want the explanation underneath the table.
- Nonparametric Statistics for the Behavioral Sciences by Sidney Siegel
- The book that brought rank-based tests into mainstream psychology. Dated in places, but it explains why the non-parametric tests were designed the way they were, which the modern textbooks tend to skip.
- Statistics Without Tears by Derek Rowntree
- A short, deliberately non-mathematical introduction to the ideas behind significance, distributions and sampling. The right starting point if the arithmetic in other books is getting in the way of the concepts.
Recommended Websites
- AQA A-level Psychology specification
- The primary source for exactly which tests are examinable and what students are expected to do with them. Section 4.2.3.3 is the one that governs test selection.
- Nature Neuroscience
- Publishes the Nieuwenhuis, Forstmann and Wagenmakers analysis of how often published studies use the wrong comparison, which is short, readable and a genuinely useful evaluation point.
- PubMed Central
- Holds the full open-access text of the Nuijten et al. survey of statistical reporting errors, along with most of the methodological literature cited here.
To cite this article please use:
Early Years TV Which Statistical Test to Use in Psychology: A Decision Guide. Available at: https://www.earlyyears.tv/which-statistical-test-to-use-psychology-decision-guide/ (Accessed: 21 September 2026).

