Mann-Whitney U Test: Formula, Worked Example and Table

Mann-Whitney U test worked example: 18 recall scores ranked together, rank totals, U = 7 and the critical value lookup

The Mann-Whitney U test is often described as a test of medians, but it is not one. Two groups with identical medians can still produce a significant Mann-Whitney result (Hart, 2001; Divine et al., 2018).

Key Takeaways

  • Use the Mann-Whitney U test when you are looking for a difference between two separate groups of people and your data are at least ordinal, so they can be put in rank order.
  • Rank every score from both groups together, add up the ranks for each group, and put each total into U = nanb + na(na + 1)/2 minus the rank total. The smaller of the two answers is U.
  • Unlike chi-squared or a t-test, a smaller number is the more impressive one here. The result is significant when U is equal to or less than the critical value in the table.

The Mann-Whitney U test is the test you reach for when two separate groups have produced scores and you want to know whether one group generally scored higher than the other. It is the independent groups partner of the Wilcoxon test, and the non-parametric alternative to the unrelated t-test. If your data are ratings, rankings, or scores that are too lopsided to trust a mean, this is usually the test you need.

Choosing it is the easy part once you know the three questions that decide every test, which are set out in our guide to choosing the right statistical test: is it a difference or a relationship, which experimental design did you use, and what is the level of measurement of your data. A test of difference with independent groups and ordinal data leads straight to Mann-Whitney.

Calculating it is where marks go missing. Tied scores, a rank total added up from the wrong group, and the backwards rule for comparing U with the critical value catch out more students than the formula itself. This guide works through all of it with a full worked example that includes ties, a second example where the result is not significant, a free calculator, and the complete critical values table. Every critical value on this page has been worked out exactly and checked against the table Pearson Edexcel prints in its specification.

What is the Mann-Whitney U test?

The Mann-Whitney U test is a non-parametric statistical test that checks whether scores in one independent group tend to be higher than scores in another. It does this by throwing away the raw scores and working with their rank order instead, then asking whether the high ranks are bunched in one group more than chance would allow.

The idea is easy to picture. Imagine lining up every participant from both groups in a single queue, from lowest score to highest. If the two conditions made no difference, the members of each group would be scattered along the queue fairly evenly. If one condition really did produce higher scores, that group would crowd towards the top end of the queue. The Mann-Whitney U test measures how lopsided the queue is, and the U statistic is the number that summarises it.

Because it only uses ranks, the test does not care how far apart the scores are, only which is bigger. A participant who recalled 9 words and one who recalled 90 would both simply be “higher” than one who recalled 8. That is exactly what makes the test safe with ordinal data and with scores that include a few extreme values, and it is also why the test is described as distribution-free: it makes no assumption that the scores follow a normal curve.

Who were Mann and Whitney?

Henry Mann was a mathematician at Ohio State University and Donald Ransom Whitney was his graduate student. In 1947 they published the paper that gave the test its name and its U statistic, and included tables of exact probabilities for small samples (Mann & Whitney, 1947). Two years earlier, the chemist Frank Wilcoxon had described a rank-sum test for two groups of equal size, in the same short paper that introduced the signed-ranks test used for repeated measures (Wilcoxon, 1945). Mann and Whitney’s contribution was to extend the approach to groups of unequal size and put it on a firm mathematical footing.

Neither pair was quite first. The statistician William Kruskal later traced essentially the same test back to a German educational psychologist, Gustav Deuchler, in 1914, and found several other independent rediscoveries along the way (Kruskal, 1957). That history explains the test’s many names. You will see it called the Mann-Whitney U test, the Wilcoxon rank-sum test, and the Wilcoxon-Mann-Whitney test, and they all give the same answer.

When should you use the Mann-Whitney U test?

Use the Mann-Whitney U test when three things are all true: you are testing for a difference rather than a correlation, your two groups contain different participants, and your data are at least ordinal. If any one of those three changes, a different test is needed.

QuestionAnswer that points to Mann-WhitneyIf the answer is different
Difference or relationship?Difference between two conditionsRelationship: Spearman’s rho
Which design?Independent groupsRepeated measures or matched pairs: Wilcoxon
Level of measurement?Ordinal (or interval data you do not trust)Nominal: chi-squared. Interval and normal: unrelated t-test

The three conditions in full

A test of difference. Your hypothesis predicts that one condition will produce different scores from another: more words recalled, higher anxiety ratings, faster reaction times. You are comparing the groups, not asking whether two measures rise and fall together. If you had a pair of scores from every participant and wanted to know whether they were related, that is a correlation, and the ordinal version is Spearman’s rho.

Independent groups. Each participant takes part in one condition only, so the two sets of scores come from different people. The groups do not need to be the same size, which is one of the practical attractions of the test: if two people drop out of one condition, you can still run it. If the same people did both conditions, or each person in one group was carefully matched to a person in the other, the scores are related and the correct test is the Wilcoxon signed-ranks test instead.

At least ordinal data. The scores must be able to go in a meaningful order. Ratings on a 1 to 10 scale, positions in a race and scores on most psychological questionnaires are ordinal: you know one is higher than another, but not that the gaps between points are equal. Mann-Whitney is also appropriate for interval data, such as time or number of words recalled, whenever those data are skewed or there are too few participants to check that they are normally distributed. What it cannot handle is nominal data, where participants are simply sorted into named categories with no order. Counts in categories need chi-squared.

Is the Mann-Whitney U test parametric or non-parametric?

The Mann-Whitney U test is non-parametric. It does not estimate population parameters such as the mean and standard deviation, and it does not assume that scores are normally distributed. Its parametric equivalent is the independent samples t-test, also called the unrelated t-test, which compares means and does require interval data with a roughly normal distribution and similar spread in both groups. Our guide to t-tests in psychology covers those criteria in full.

Why use Mann-Whitney instead of a t-test?

Use Mann-Whitney instead of a t-test when your data are ordinal, when they are clearly skewed, or when one or two extreme scores would drag a mean around. A t-test works with the actual size of every score, so a single outlier can inflate the mean and the standard deviation and change the result. Mann-Whitney only knows that the outlier is the highest score, so it barely notices how extreme it is.

Students often assume the trade-off is a big loss of power, because ranking throws information away. The mathematics says otherwise. Hodges and Lehmann (1956) showed that even when the data really are perfectly normal, the rank-sum test is about 95% as efficient as the t-test, and that across all continuous distributions its efficiency never falls below about 86%. With heavy-tailed or skewed data it can be far more powerful than the t-test. In practice that means you give up very little by choosing Mann-Whitney, and sometimes you gain a great deal.

The Mann-Whitney U formula explained

The Mann-Whitney U formula calculates a U value for each group from that group’s size and its total of ranks, and the smaller of the two values is the test statistic:

Ua = nanb + na(na + 1) / 2 − ΣRa
Ub = nanb + nb(nb + 1) / 2 − ΣRb
U is the smaller of Ua and Ub

Every symbol has a plain meaning:

  • na and nb are the number of participants in group A and group B.
  • ΣRa and ΣRb are the rank totals: the sum of the ranks given to the scores in each group, after every score from both groups has been ranked together. The Greek capital sigma, Σ, just means “add them all up”.
  • nanb means na multiplied by nb.

This is the form printed in the Pearson Edexcel specification, and it is the one you will meet most often in UK textbooks (Pearson Edexcel, 2026). You only strictly need to calculate one of the two U values, because of a useful check: Ua and Ub always add up to na multiplied by nb. Work out one, subtract it from nanb, and you have the other. Calculating both and confirming they add up is the quickest way to catch an arithmetic slip.

There is a second check worth building into every calculation. The two rank totals together must equal N(N + 1)/2, where N is the total number of participants. With 18 participants, the ranks 1 to 18 add up to 18 × 19 / 2 = 171, so ΣRa + ΣRb must come to 171. If it does not, a rank has been missed or a tie has been averaged incorrectly.

What U actually counts

U has a meaning that the formula hides. If you took every possible pairing of one person from group A with one person from group B, U counts how many times the person from the lower-scoring group came out on top, with a tie counting as half. With 8 people in one group and 10 in the other there are 80 possible pairings, which is why Ua and Ub always sum to 80: every pairing is won by one side or the other.

So a U of zero means the two groups do not overlap at all: every single score in one group beats every single score in the other. A U close to half of nanb means the groups are thoroughly mixed, which is what the null hypothesis predicts. That is why small values of U are the significant ones, and why the comparison with the critical value runs the opposite way to chi-squared and the t-test. It also gives you a second method of calculation. Some textbooks teach this counting method instead of the formula, and both give the same U, as the worked example below shows.

Mann-Whitney U test worked example: 18 recall scores ranked together, rank totals, U = 7 and the critical value lookup
Every number in the diagram comes from the worked example below. Notice that the group with the lower scores ends up with the smaller rank total but the larger U, which is why you always take the smaller of the two U values.

How to calculate the Mann-Whitney U test: a worked example

To calculate the Mann-Whitney U test, rank all the scores from both groups together, add up the ranks for each group, put those totals into the formula, take the smaller U, and compare it with the critical value. The example below walks through each step with real numbers, including tied scores, which is where most mistakes happen.

The study is a small replication of a classic memory experiment. Baddeley (1966) found that lists of words that sound alike, such as man, cap, can and mad, are harder to hold in short-term memory than lists of words that sound different. This fits the idea that short-term memory relies heavily on an acoustic code, a key part of the multi-store model of memory. The data below are invented for teaching purposes, but the design is exactly what a student practical would use.

Eighteen participants were randomly allocated to one of two conditions. Eight heard a list of ten acoustically similar words (group A) and ten heard a list of ten acoustically dissimilar words (group B). Each then wrote down as many words as they could in the correct order, and the score is the number recalled correctly out of ten. The groups are different sizes on purpose, because that is something Mann-Whitney handles without any fuss.

Step 1: State the hypotheses

Because previous research predicts which way the difference should go, the hypothesis is directional, which means a one-tailed test.

  • Experimental hypothesis: Participants who hear acoustically similar words will recall fewer words in the correct order than participants who hear acoustically dissimilar words.
  • Null hypothesis: There will be no difference in the number of words recalled in the correct order between participants who hear acoustically similar words and participants who hear acoustically dissimilar words.

Step 2: Set out the scores

GroupScores (words recalled out of 10)nMedian
A: acoustically similar3, 4, 4, 5, 5, 6, 6, 785
B: acoustically dissimilar5, 6, 7, 7, 8, 8, 8, 9, 9, 10108

Before doing any arithmetic, look at the data. Group B’s scores are clearly higher on the whole, but the two groups overlap between 5 and 7. The test will tell us whether that overlap is small enough for the difference to count as significant.

Step 3: Rank all the scores together

Put all eighteen scores into one list, lowest first, and give the lowest score rank 1. Where two or more scores are tied, give each of them the average of the positions they occupy. Group membership plays no part in the ranking at all: you rank the pooled scores and only split them back into groups afterwards.

ScoreHow many (A / B)Positions occupiedRank given to each
31 / 011
42 / 02, 32.5
52 / 14, 5, 65
62 / 17, 8, 98
71 / 210, 11, 1211
80 / 313, 14, 1514
90 / 216, 1716.5
100 / 11818

Take the three scores of 5 as an example. They fill positions 4, 5 and 6 in the pooled list, so each is given (4 + 5 + 6) / 3 = 5. The next score, the first 6, then carries on from position 7, not from 6. Forgetting to skip on after a tie is the single most common error in the whole calculation.

Step 4: Add up the ranks for each group

  • Group A ranks: 1 + 2.5 + 2.5 + 5 + 5 + 8 + 8 + 11 = 43
  • Group B ranks: 5 + 8 + 11 + 11 + 14 + 14 + 14 + 16.5 + 16.5 + 18 = 128

Check: 43 + 128 = 171, and 18 × 19 / 2 = 171. The totals agree, so no rank has been lost.

Step 5: Calculate U for each group

Here na = 8, nb = 10, so nanb = 80.

  • Ua = 80 + (8 × 9) / 2 − 43 = 80 + 36 − 43 = 73
  • Ub = 80 + (10 × 11) / 2 − 128 = 80 + 55 − 128 = 7

Check: 73 + 7 = 80, which is nanb. Both values are right.

Step 6: Take the smaller value as U

U is the smaller of 73 and 7, so U = 7. Notice that it came from group B, the group with the higher scores. That surprises people, but it follows from what U counts: 7 is the number of times, out of the 80 possible pairings, that a participant from the similar-words group beat a participant from the dissimilar-words group.

You can confirm this with the counting method. For each score in group A, count how many group B scores are lower, adding a half for each tie. The score of 3 beats no one and nor do the two 4s. Each 5 ties with one B score, giving 0.5 each. Each 6 beats the B score of 5 and ties with the B score of 6, giving 1.5 each. The 7 beats the B scores of 5 and 6 and ties with two 7s, giving 3. Adding up: 0 + 0 + 0 + 0.5 + 0.5 + 1.5 + 1.5 + 3 = 7. Same answer.

Step 7: Find the critical value

You need three things to find the critical value: the two group sizes, whether the test is one-tailed or two-tailed, and the significance level. Here that is na = 8 and nb = 10, one-tailed, at p ≤ 0.05. In the one-tailed table below, the critical value for 8 and 10 is 20.

Step 8: Compare and state the conclusion

U = 7 is less than the critical value of 20, so the result is significant at p ≤ 0.05. It is in fact below the one-tailed critical value of 13 for p ≤ 0.01 as well, so the result is significant at the stricter level too. The null hypothesis is rejected and the experimental hypothesis is accepted: participants who heard acoustically similar words recalled significantly fewer words in the correct order (median = 5) than participants who heard acoustically dissimilar words (median = 8).

The final step, which examiners reward, is to say what this means psychologically. The result supports the claim that information in short-term memory is coded acoustically, because words that sound alike become confused with one another. Our article on working memory and short-term memory explains how Baddeley’s later model built this finding into the phonological loop.

Mann-Whitney U test calculator and critical values table

Type or paste the scores for your two groups and the calculator ranks them, handles any ties, works out Ua, Ub and U, and gives the verdict against the correct critical value. It shows every rank, so you can check your own working line by line rather than just copying the answer. The empty boxes show the worked example’s scores as a guide. Both critical values tables sit underneath.

Mann-Whitney U Test Calculator

Type or paste the scores for each group, separated by commas or spaces. Independent groups only. Groups can be different sizes.

Critical values of U. n1 and n2 are the sizes of the two groups, and it does not matter which is which. Your calculated U must be equal to or less than the critical value for the result to be significant.

Table 1: p ≤ 0.05 two-tailed (p ≤ 0.025 one-tailed)
n1 \ n2567891011121314151617181920
52356789111213141517181920
63568101113141617192122242527
756810121416182022242628303234
8681013151719222426293134363841
97101215172023262831343739424548
108111417202326293336394245485255
119131619232630333740444751555862
1211141822262933374145495357616569
1312162024283337414550545963677276
1413172226313640455055596469747883
1514192429343944495459647075808590
1615212631374247535964707581869298
17172228343945515763697581879399105
181824303642485561677480869399106112
1919253238455258657278859299106113119
20202734414855626976839098105112119127
Table 2: p ≤ 0.05 one-tailed (p ≤ 0.10 two-tailed)
n1 \ n2567891011121314151617181920
5456891112131516181920222325
657810121416171921232526283032
7681113151719212426283033353739
88101315182023262831333639414447
99121518212427303336394245485154
1011141720242731343741444851555862
1112161923273134384246505457616569
1213172126303438424751556064687277
1315192428333742475156616570758084
1416212631364146515661667177828792
15182328333944505561667277838894100
161925303642485460657177838995101107
1720263339455157647077838996102109115
18222835414855616875828895102109116123
192330374451586572808794101109116123130
2025323947546269778492100107115123130138

The calculator uses exact critical values for any combination of group sizes from 2 to 20, at both p ≤ 0.05 and p ≤ 0.01. When either group has more than 20 participants the tables run out, and it switches to the normal approximation explained further down this page.

A second worked example: when the result is not significant

A non-significant result is reported in exactly the same way as a significant one, and it is just as common in real research. This shorter example uses ordinal data, a non-directional hypothesis and groups of 6 and 7.

A researcher wanted to know whether the style of a police interview affects how confident eyewitnesses feel about their memory of a filmed event. Six participants were interviewed in a standard way (group A) and seven with a more structured memory-prompting technique (group B). Afterwards each rated their confidence on a scale of 1 to 10. A confidence rating is ordinal data, since there is no reason to think the gap between 3 and 4 means the same as the gap between 8 and 9, so a t-test would be inappropriate. The data are invented.

GroupConfidence ratings (1 to 10)RanksRank total
A: standard interview (n = 6)4, 6, 7, 7, 8, 91, 3.5, 5.5, 5.5, 7.5, 1033
B: structured interview (n = 7)5, 6, 8, 9, 9, 10, 102, 3.5, 7.5, 10, 10, 12.5, 12.558

The check first: 33 + 58 = 91, and 13 × 14 / 2 = 91. Then the formula, with nanb = 42:

  • Ua = 42 + (6 × 7) / 2 − 33 = 42 + 21 − 33 = 30
  • Ub = 42 + (7 × 8) / 2 − 58 = 42 + 28 − 58 = 12

30 + 12 = 42, so the arithmetic is sound, and U = 12. For groups of 6 and 7, two-tailed at p ≤ 0.05, the critical value is 6. U = 12 is greater than 6, so the result is not significant. The null hypothesis is retained: there was no significant difference in confidence ratings between participants given a standard interview (median = 7) and those given a structured interview (median = 9), U = 12, p > 0.05, two-tailed.

The medians differ by two whole points, which looks like a sizeable gap. Yet with only thirteen participants and plenty of overlap between the groups, that gap is well within what chance alone could produce. That does not prove the interview style makes no difference. It means this study did not find evidence that it does, and a larger sample might. Failing to detect a real effect is a Type II error, which our guide to statistical significance, p-values and errors explains alongside its opposite.

How to read the Mann-Whitney U critical values table

To read the Mann-Whitney U critical values table, pick the table for your significance level and type of hypothesis, find the row for one group size and the column for the other, and read off the number where they meet. Your U must be equal to or less than that number for the result to be significant.

Unlike most tests, the Mann-Whitney table has two dimensions, because the critical value depends on both group sizes rather than on a single N or a degrees of freedom figure. It does not matter which group you treat as n1 and which as n2: the table is symmetrical, so the value for 8 and 10 is the same as the value for 10 and 8.

Why U must be equal to or less than the critical value

U must be equal to or less than the critical value because a small U means little overlap between the groups, which is what a real difference looks like. This is the rule students most often get backwards, because chi-squared, the t-tests and Spearman’s rho all work the other way: for those tests the calculated value must be equal to or greater than the critical value. The tests that share the Mann-Whitney rule are the other two that count up the “losing” side, the sign test and the Wilcoxon test. A well-known memory aid: if there is an R in the name of the test, as in Spearman’s rho, Pearson’s r, chi-squared and the related and unrelated t-tests, the calculated value must be gReater than or equal to the critical value. The sign test, Wilcoxon and Mann-Whitney have no R, so their calculated value must be equal to or less. Examiners often state the rule in the question, but not always, so learn it.

One-tailed or two-tailed?

Use the one-tailed values when your hypothesis is directional, predicting which group will score higher, and the two-tailed values when it simply predicts a difference. The same table serves both, because a one-tailed probability of 0.025 corresponds to a two-tailed probability of 0.05. That is why each table heading carries two numbers. There is one extra check for a one-tailed test: the difference must actually be in the direction you predicted. If you predicted that group A would score higher and group B did, the result does not support your hypothesis however small U is.

What the dashes in the table mean

With very small groups, some cells have no critical value at all, and printed tables show a dash. A dash means that even a U of 0, with the groups completely separated, would not be significant at that level. Two groups of three, for example, can never produce a two-tailed result significant at p ≤ 0.05, because there are only 20 possible ways to split six ranks between them and the most extreme split happens by chance 2 times in 20, a probability of 0.10. If your groups are this small, the fix is more participants, not a different table.

A note on the Edexcel table

The critical values on this page were calculated exactly from the distribution of U, and they match the table in the Pearson Edexcel specification in every cell but two. For groups of 14 and 17 at p ≤ 0.05 two-tailed, Edexcel prints 67 where the exact value is 69. For groups of 13 and 19 at p ≤ 0.01 two-tailed, it prints 56 where the exact value is 57. In both cases the Edexcel figure is slightly stricter, so it could never make a result look significant when it is not. In an exam, always use the table printed on the paper or in your board’s formula booklet: that is what the mark scheme will follow. The differences only matter if a result falls exactly on 68 or 69, or exactly on 57, with those group sizes.

How to deal with tied ranks

Deal with tied scores by giving each of them the average of the rank positions they would have taken up, then carrying on from the next free position. Ties between the two groups and ties within one group are handled in exactly the same way, because the ranking is always done on the pooled scores.

  1. Write out every score from both groups in one list, lowest to highest.
  2. Number the positions 1, 2, 3 and so on down the list, ignoring ties for now.
  3. For any run of equal scores, add up the positions they occupy and divide by how many there are.
  4. Give that average to every score in the run.
  5. Carry on with the next score at the next position number.

The rank total check, N(N + 1)/2, still works perfectly with ties, because averaging never changes the total. That makes it the ideal way to catch a tie that has been ranked wrongly.

Lots of ties are a sign that the measure is too coarse. A three-point rating scale with twenty participants will produce huge blocks of tied ranks, and the test loses much of its ability to spot a difference. At A level you just rank the ties and carry on. In published research, the large-sample version of the test includes a small correction for ties, which statistics software applies automatically.

Large samples: the z approximation

When either group has more than 20 participants, the critical values tables run out and U is converted into a z score instead. As samples grow, the distribution of U gets closer and closer to a normal curve, so it can be compared with the familiar critical values of z (Mann & Whitney, 1947).

z = (U − nanb / 2) ÷ √[ nanb(na + nb + 1) / 12 ]

The top line compares U with the value it would have on average if the null hypothesis were true, which is half of nanb. The bottom line is the standard deviation of U under the null hypothesis, playing the same role as the standard error does for a mean. Ignoring the minus sign, the result is significant two-tailed at p ≤ 0.05 if z is 1.96 or more, and at p ≤ 0.01 if it is 2.58 or more. For a one-tailed test the cut-offs are 1.645 and 2.33.

Applied to the first worked example, the mean of U is 40 and its standard deviation is the square root of 80 × 19 / 12, which is about 11.25. So z = (7 − 40) / 11.25 = −2.93. That comfortably passes both cut-offs, which agrees with the exact table. With groups as small as 8 and 10, though, the exact table is the right tool, and the approximation is only a cross-check.

How to write up a Mann-Whitney U result

Write up a Mann-Whitney result by giving the direction of the difference, the medians of both groups, the value of U, the group sizes, the significance level, and whether the test was one-tailed or two-tailed, then say whether the null hypothesis is rejected or retained. Medians are the right descriptive statistic to report here, because the test is based on ranks rather than means.

For the first worked example, an A level write-up would read:

A Mann-Whitney U test was used because the design was independent groups, the hypothesis predicted a difference, and the data were treated as ordinal. Participants who heard acoustically similar words recalled significantly fewer words in the correct order (median = 5, n = 8) than participants who heard acoustically dissimilar words (median = 8, n = 10), U = 7, p ≤ 0.01, one-tailed. As the calculated value of 7 is less than the critical value of 13, the null hypothesis is rejected.

In a university report you would normally add the exact p value and an effect size, which software gives you. When a question asks you to justify the choice of test, give all three reasons, difference, independent groups and level of measurement, and link each one to the study in the question rather than to the test in general.

What does the Mann-Whitney U test actually show?

The Mann-Whitney U test shows whether a randomly chosen score from one group is more likely to be higher than a randomly chosen score from the other group than the reverse. That is subtly different from asking whether the two medians differ, and the difference matters once you move past A level.

If the two groups have the same shape of distribution and the same spread, and one is simply shifted up the scale compared with the other, then a significant Mann-Whitney result does mean the medians differ. That shift assumption is what textbooks have in mind when they call Mann-Whitney a test of medians. When the shapes or spreads differ, the link breaks. Hart (2001) showed that groups with equal medians can give a significant result when their spreads differ, and Divine et al. (2018) showed that the test can fail badly when used as a test of medians. Fay and Proschan (2010) set out exactly which assumptions justify which interpretation.

For exam purposes, stating that the test looks at whether scores in one group tend to be higher than in the other, and reporting the medians as the descriptive statistic, is accurate and will gain credit. The deeper point is worth knowing because it explains why a researcher should always look at the distribution of the data, not just the verdict. A quick look at a dot plot of both groups side by side shows at once whether one group is simply shifted or whether something more complicated is going on.

Effect size for the Mann-Whitney U test

The simplest effect size for the Mann-Whitney U test is the rank-biserial correlation, calculated as r = 1 − 2U / (nanb). It runs from 0, meaning the groups are completely mixed, to 1, meaning they do not overlap at all (Wendt, 1972).

For the first worked example, r = 1 − (2 × 7) / 80 = 1 − 0.175 = 0.825, a very large effect. Kerby (2014) showed that the same number can be read in an intuitive way, as the proportion of pairings that favour the higher group minus the proportion that favour the lower group. Out of the 80 pairings, 73 favoured group B and 7 favoured group A, so 91.25% minus 8.75% gives 0.825. Put simply, a participant who heard dissimilar words beat a participant who heard similar words in about nine pairings out of ten.

Effect size matters because significance depends heavily on sample size. A tiny difference can be significant with hundreds of participants, and a big difference can miss significance with a dozen, as the second worked example showed. The effect size tells you how big the difference is, whatever the sample size. It is not on the A level specifications for Mann-Whitney, but it is expected in almost every university report. Another common option is r = z / √N, which software such as SPSS makes easy to calculate from its reported z.

Mann-Whitney vs the t-test, Wilcoxon and other tests

Mann-Whitney is one of several inferential tests on the A level specifications, and each covers one combination of purpose, design and level of measurement. The table below shows how Mann-Whitney relates to the tests it is most often confused with.

TestPurposeDesignDataSignificant when
Mann-Whitney UDifferenceIndependent groupsOrdinalU ≤ critical value
Wilcoxon signed-ranksDifferenceRepeated measures or matched pairsOrdinalT ≤ critical value
Unrelated t-testDifferenceIndependent groupsInterval, normalt ≥ critical value
Sign testDifferenceRepeated measures or matched pairsNominalS ≤ critical value
Chi-squaredDifference or associationIndependent groupsNominalχ² ≥ critical value
Spearman’s rhoCorrelationPairs of scoresOrdinalrho ≥ critical value

Mann-Whitney vs Wilcoxon

Mann-Whitney and the Wilcoxon signed-ranks test both work with ranks and both test for a difference, but Mann-Whitney is for two separate groups and Wilcoxon is for the same people tested twice or for matched pairs. The confusion is made worse by the names, since the Wilcoxon rank-sum test is Mann-Whitney under another name. At A level, “Wilcoxon” always means the signed-ranks test for related designs, and the design of the study is what decides between them.

Mann-Whitney vs chi-squared

Mann-Whitney and chi-squared both suit independent groups, so the level of measurement decides between them. If each participant has a score that can be put in order, use Mann-Whitney. If each participant has simply been counted into a named category, such as helped or did not help, use chi-squared.

Mann-Whitney vs Kruskal-Wallis

The Kruskal-Wallis test extends the same ranking logic to three or more independent groups. It is not on the A level specifications, but you will meet it at university. With exactly two groups, Kruskal-Wallis is equivalent to a two-tailed Mann-Whitney test.

Common Mann-Whitney mistakes in exams

The most common Mann-Whitney mistakes are ranking each group separately, mishandling ties, and comparing U with the critical value the wrong way round. All of them are avoidable with the two arithmetic checks described above.

  • Ranking the groups separately. Both groups must be ranked together as one list. Ranking each group from 1 upwards makes the test meaningless.
  • Not skipping positions after a tie. Three tied scores in positions 4, 5 and 6 all get rank 5, and the next score gets rank 7.
  • Taking the larger U. U is always the smaller of Ua and Ub.
  • Reading the comparison backwards. For Mann-Whitney, U must be equal to or less than the critical value.
  • Using the wrong table for the hypothesis. A directional hypothesis needs the one-tailed values.
  • Choosing Mann-Whitney for a repeated measures design. If the same people did both conditions, the test is Wilcoxon.
  • Justifying the choice with “the data are non-parametric”. Data are not parametric or non-parametric; tests are. Say that the data are ordinal, or that interval data are skewed.
  • Reporting means. Report the median of each group alongside a Mann-Whitney result.

What each exam board requires

All three main A level boards include the Mann-Whitney U test, but they differ in whether you must be able to calculate it or only to choose and interpret it. Check your own specification, because the difference decides how you revise.

BoardWhat the specification saysWhat that means for you
AQA (7182)Students must know when to use Mann-Whitney among the tests listed, and use statistical tables and critical values. Calculation is specified only for the sign test.Justify choosing it, and compare a given U with a critical value. Knowing the calculation helps you understand what you are comparing.
Pearson Edexcel (9PS0)Mann-Whitney U is one of the two non-parametric tests of difference, used in the cognitive practical, and the formula and critical values tables are printed in the specification.Be ready to calculate U from raw data and interpret it, using the printed formula and tables.
OCR (H567)The criteria for using, and the use of, the Mann-Whitney U test are listed among the non-parametric tests in the research methods content.Know when and why to use it, and how to interpret U against a critical value.

The wording above comes from each board’s current specification (AQA, 2015; Pearson Edexcel, 2026; OCR, 2026). Whichever board you sit, the reasoning behind the choice is examined everywhere, and a solid revision method such as practising past questions under timed conditions pays off more than rereading notes.

Where the Mann-Whitney U test is used in psychology

The Mann-Whitney U test is used wherever psychologists compare two separate groups on a measure that is ordinal, skewed, or collected from small samples. That covers a large share of real research, because many psychological measures are ratings rather than true measurements.

  • Clinical and health psychology: comparing symptom ratings between a treatment group and a control group, where scores are often skewed because most people cluster at the mild end.
  • Developmental psychology: comparing two age groups on a task with a small number of possible scores, such as how many conservation problems children solve.
  • Cognitive psychology: comparing reaction times between conditions, which are almost always skewed by a few very slow responses.
  • Student practicals: any independent groups experiment with a rating scale or a small sample, which is why it appears so often in coursework-style questions.

Mann-Whitney only works on data that can be put in order, so it belongs with quantitative analysis. Qualitative material such as interview transcripts would first need to be coded into scores before any test could be applied.

Conclusion

The Mann-Whitney U test compares two independent groups by ranking all the scores together and asking whether the high ranks fall mainly in one group. Choose it for a test of difference with independent groups and ordinal data, or with interval data that are skewed. Rank the pooled scores and average any ties, add up the ranks for each group, calculate Ua and Ub, and take the smaller. Check that the rank totals add to N(N + 1)/2 and that the two U values add to nanb. The result is significant when U is equal to or less than the critical value, the opposite way round from chi-squared and the t-test. Report the medians, and remember that the test is really about which group tends to score higher, not strictly about the medians themselves.

Frequently Asked Questions

What is the Mann-Whitney U test used for?

It is used to decide whether two separate groups of people differ on some measure, when that measure can be put in rank order but cannot be trusted to behave like a normal, evenly spaced scale. Typical examples are ratings, questionnaire scores, and small samples of recall or reaction time data.

Is Mann-Whitney parametric or non-parametric?

Non-parametric. It works on ranks, so it makes no assumption that scores follow a normal distribution and does not estimate a population mean or standard deviation. The unrelated t-test is its parametric counterpart.

What is the difference between the Mann-Whitney U test and a t-test?

Both compare two independent groups. The unrelated t-test compares means and needs interval data that are roughly normal with similar variance in each group. Mann-Whitney compares rank positions, so it copes with ordinal data, skew and outliers, at the cost of a small loss of power when the data really are normal.

Are Mann-Whitney and Wilcoxon the same?

Mann-Whitney is identical to the Wilcoxon rank-sum test, which is why software sometimes reports both names. It is not the same as the Wilcoxon signed-ranks test, which is for related designs where each participant produces two scores. At A level, “Wilcoxon” refers to the signed-ranks version.

What level of measurement does the Mann-Whitney U test need?

Ordinal or higher. Any data that can be placed in a meaningful order will do, including interval and ratio data. Nominal data, which only sort people into unordered categories, cannot be ranked and need chi-squared instead.

How do you calculate the Mann-Whitney U test by hand?

  1. Pool both groups and rank every score, averaging tied positions.
  2. Total the ranks separately for group A and group B.
  3. Substitute each total into its U formula.
  4. Keep whichever U is smaller.
  5. Look up the critical value for your two group sizes and hypothesis; significance needs U at or below it.

When should you use Mann-Whitney instead of a t-test?

Choose Mann-Whitney whenever the scores are ratings or rankings, whenever a histogram shows obvious skew, or when a couple of extreme values would distort a mean. With small samples it is also the safer option, since normality is almost impossible to check with only a handful of participants per group.

How do you calculate effect size for a Mann-Whitney U test?

The easiest measure is the rank-biserial correlation: double U, divide by the product of the two group sizes, and subtract the result from 1. Values near 0 mean heavy overlap between the groups and values near 1 mean almost none. An alternative is to divide the z score from software by the square root of the total sample size.

What is the difference between Mann-Whitney and Kruskal-Wallis?

Mann-Whitney handles exactly two independent groups, while Kruskal-Wallis handles three or more. Both rank the pooled scores. Running several Mann-Whitney tests instead of one Kruskal-Wallis inflates the chance of a false positive, so a researcher with three groups should start with Kruskal-Wallis.

References

AQA. (2015). AS and A-level Psychology specification (7181, 7182). Assessment and Qualifications Alliance.

Baddeley, A. D. (1966). Short-term memory for word sequences as a function of acoustic, semantic and formal similarity. Quarterly Journal of Experimental Psychology, 18(4), 362-365.

Divine, G. W., Norton, H. J., Barón, A. E., & Juarez-Colunga, E. (2018). The Wilcoxon-Mann-Whitney procedure fails as a test of medians. The American Statistician, 72(3), 278-286.

Fay, M. P., & Proschan, M. A. (2010). Wilcoxon-Mann-Whitney or t-test? On assumptions for hypothesis tests and multiple interpretations of decision rules. Statistics Surveys, 4, 1-39.

Hart, A. (2001). Mann-Whitney test is not just a test of medians: Differences in spread can be important. BMJ, 323(7309), 391-393.

Hodges, J. L., & Lehmann, E. L. (1956). The efficiency of some nonparametric competitors of the t-test. The Annals of Mathematical Statistics, 27(2), 324-335.

Kerby, D. S. (2014). The simple difference formula: An approach to teaching nonparametric correlation. Comprehensive Psychology, 3, Article 11.IT.3.1.

Kruskal, W. H. (1957). Historical notes on the Wilcoxon unpaired two-sample test. Journal of the American Statistical Association, 52(279), 356-360.

Mann, H. B., & Whitney, D. R. (1947). On a test of whether one of two random variables is stochastically larger than the other. The Annals of Mathematical Statistics, 18(1), 50-60.

OCR. (2026). A Level Psychology H567 specification (Version 1.5). Oxford Cambridge and RSA Examinations.

Pearson Edexcel. (2026). Specification: Pearson Edexcel Level 3 Advanced GCE in Psychology (9PS0) (Issue 4). Pearson Education Limited.

Wendt, H. W. (1972). Dealing with a common problem in social science: A simplified rank-biserial coefficient of correlation based on the U statistic. European Journal of Social Psychology, 2(4), 463-465.

Wilcoxon, F. (1945). Individual comparisons by ranking methods. Biometrics Bulletin, 1(6), 80-83.

Further Reading and Research

Recommended Articles

Suggested Books

  • Research Methods and Statistics in Psychology by Hugh Coolican
    • The standard UK text for psychology students moving from A level to degree study. It explains every A level test, including Mann-Whitney, with worked examples and clear guidance on choosing between them.
  • Discovering Statistics Using IBM SPSS Statistics by Andy Field
    • The friendliest serious statistics book in print. Its chapter on non-parametric tests shows how Mann-Whitney is run and reported in software, including effect sizes.
  • Nonparametric Statistics for the Behavioral Sciences by Sidney Siegel and N. John Castellan
    • The classic reference for rank-based tests in psychology, with the logic, assumptions and tables for Mann-Whitney and its relatives set out in full.

Recommended Websites

  • Project Euclid: Mann and Whitney (1947)
    • The original paper that introduced the U statistic, free to read, including the first tables of exact probabilities for small samples.
  • Pearson Edexcel Psychology qualification pages
    • The Edexcel specification prints the Mann-Whitney formula and four full critical values tables in its appendix, alongside the Wilcoxon process.
  • AQA A-level Psychology specification
    • Section 4.2.3 sets out which inferential tests students must know and what they must be able to do with each.

Kathy Brodie

Kathy Brodie is an Early Years Professional, Trainer and Author of multiple books on Early Years Education and Child Development. She is the founder of Early Years TV and the Early Years Summit.

Kathy’s Author Profile
Kathy Brodie

To cite this article please use:

Early Years TV Mann-Whitney U Test: Formula, Worked Example and Table. Available at: https://www.earlyyears.tv/mann-whitney-u-test-formula-worked-example-table/ (Accessed: 2 October 2026).