Wilcoxon Signed-Rank Test: Formula, Worked Example and Table

Wilcoxon signed-rank test worked example: differences ranked by size, T = 6.5 with N = 11 against a critical value of 10

In 1945 the chemist Frank Wilcoxon introduced the Wilcoxon signed-rank test in a four-page paper and tried it on Darwin’s maize plants. It reached the same verdict as Fisher’s t-test without calculating a single mean.

Key Takeaways

  • Use the Wilcoxon signed-rank test when you are looking for a difference between two related conditions, from a repeated measures or matched pairs design, and your data are at least ordinal.
  • Work out each participant’s difference, drop any zeros, rank the differences by size while ignoring the sign, then add up the positive ranks and the negative ranks separately. T is the smaller of the two totals.
  • A small T is the impressive result. It is significant when T is equal to or less than the critical value for N, the number of pairs that actually differ.

The Wilcoxon signed-rank test is the test you reach for when the same people have been measured twice, or when carefully matched pairs have each done one condition, and you want to know whether one condition generally produced higher scores than the other. It is the related-design partner of the Mann-Whitney U test and the non-parametric alternative to the related t-test. If your data are ratings, rankings, or scores too lopsided to trust a mean, and every score has a partner, this is usually the test you need.

Picking it takes three questions, which are set out in our guide to choosing the right statistical test: are you looking for a difference or a relationship, which experimental design did you use, and what is the level of measurement of your data? A difference, a related design and ordinal data lead straight to Wilcoxon. If your data only tell you which condition was better, not by how much, the simpler sign test is the right choice instead.

Calculating it is where marks go missing. Zero differences left in, ranks given to the raw scores instead of the differences, and the backwards rule for comparing T with the critical value catch out more students than the arithmetic. This guide works through all of it: a full worked example with a zero and tied ranks, a second example where the result is not significant, a free calculator, and a complete critical values table from N = 5 to 50. Every critical value on this page has been calculated exactly and checked against the tables printed by Pearson Edexcel and used in AQA exam questions, and the method has been checked against Wilcoxon’s own 1945 examples.

What is the Wilcoxon signed-rank test?

The Wilcoxon signed-rank test is a non-parametric statistical test that checks whether scores in one condition tend to be higher than scores in another, when the scores come in related pairs. It works on the difference between each pair of scores. It ranks those differences by size, gives each rank the sign of its difference, and then asks whether the big differences mostly point one way.

The logic is easy to picture. Suppose ten people each learn a word list in silence and another with music playing. For each person, take their music score away from their silence score. If music made no difference, some people would do a little better with music and some a little worse, and the large differences would be just as likely to fall on either side. If silence really helps, most of the differences will be positive, and crucially the big differences will be positive too, with only a few small negative ones. The test measures exactly that imbalance.

Because it works with ranks rather than the raw differences, the test does not need the differences to follow a normal curve and is not thrown by one person with a freakishly large change. That is what makes it safe with ordinal data such as ratings, and it is why it is described as distribution-free. But unlike the sign test, it does not throw away the size of each difference completely. A difference of 7 words outranks a difference of 1 word, and that extra information is what makes Wilcoxon the more powerful of the two.

Who was Frank Wilcoxon?

Frank Wilcoxon was a chemist at the American Cyanamid Company, working on experiments such as fly sprays and seed treatments rather than on psychology. His paper Individual Comparisons by Ranking Methods, published in Biometrics Bulletin in December 1945, runs from page 80 to page 83 and introduced two tests at once (Wilcoxon, 1945). One was a rank-sum test for two separate groups, which became the Mann-Whitney U test after Mann and Whitney extended it in 1947. The other was the signed-rank test for paired data, and that is the one that still carries his name alone.

His stated aim was modest. He wanted ranking methods that would give “a rapid approximate idea of the significance of the differences” (Wilcoxon, 1945, p. 80), without the heavy arithmetic of a t-test or analysis of variance done by hand. More than eighty years later, that speed is exactly why the test suits an exam hall.

What does “signed-rank” mean?

“Signed-rank” means that each difference is given a rank based on its size, and that rank then carries the plus or minus sign of the original difference. A difference of −3 that is the fifth smallest difference becomes the signed rank −5, for example. The test statistic, T, is simply the total of the ranks on the less common side.

The test goes by several names. AQA exam questions, the OCR specification and the Pearson Edexcel specification all call it the Wilcoxon signed ranks test, revision guides often shorten it to “the Wilcoxon test”, and textbooks sometimes say Wilcoxon matched-pairs signed-ranks test or Wilcoxon T test. They are all the same procedure.

When should you use the Wilcoxon signed-rank test?

Use the Wilcoxon signed-rank test when three things are all true: you are testing for a difference, the two sets of scores are related because they come from a repeated measures or matched pairs design, and the data are at least ordinal. If any one of those changes, a different test is needed.

QuestionAnswer that leads to WilcoxonIf the answer is different
Difference or relationship?Difference between two conditionsRelationship: Spearman’s rho or Pearson’s r
Which design?Repeated measures or matched pairsIndependent groups: Mann-Whitney U
Level of measurement?Ordinal, or interval data you do not trust to be normalNominal: sign test. Normal interval data: related t-test
The three questions that lead to the Wilcoxon signed-rank test.

The three conditions in full

A test of difference. The hypothesis predicts that scores will be higher in one condition than the other, such as “participants recall more words in silence than with music playing”. A hypothesis about scores going up together, such as “the more hours of sleep, the higher the recall score”, is a correlation and needs a different test.

A related design. Every score in condition A has a partner in condition B. In a repeated measures design the partner is the same person tested again. In a matched pairs design the partner is a different person chosen to be similar, for example on age, IQ or baseline score. What matters is that the pairing is real and fixed before the data are collected, because the whole test is built on the difference within each pair. If you could shuffle the order of the condition B scores without changing anything, the design is not related and Wilcoxon is the wrong test.

At least ordinal data. The differences need to be rankable by size. That is automatically true of interval and ratio data such as reaction times and words recalled, and true of most rating scales. Strictly, the test also assumes that a difference of 3 is genuinely bigger than a difference of 2 wherever on the scale it occurs, which is why some statisticians prefer the sign test for rough ratings. At A level, “ordinal or above” is the accepted rule.

Is the Wilcoxon signed-rank test parametric or non-parametric?

The Wilcoxon signed-rank test is non-parametric. A parametric test such as the related t-test assumes that the differences come from a normally distributed population and uses the mean and standard deviation of the differences. The Wilcoxon test makes no assumption about normality and uses only the ranks of the differences.

Non-parametric does not mean assumption-free. The pairs must be independent of one another, so one participant’s result cannot influence another’s. And when the test is used to make a claim about the median difference specifically, the differences are assumed to be roughly symmetrical around that median (Conover, 1999). For the A level question “did one condition generally produce higher scores than the other?”, those conditions are easily met.

Why use Wilcoxon instead of the related t-test?

Use Wilcoxon instead of the related t-test when the data are ordinal, when the differences are clearly skewed, or when there are extreme values that would drag the mean around. With small samples, which is what most psychology practicals produce, it is hard to show that differences are normally distributed at all, and Wilcoxon avoids having to try.

The price of that safety is smaller than many students assume. Hodges and Lehmann (1956) showed that, for large samples, the efficiency of the Wilcoxon test relative to the t-test never falls below 0.864, whatever the shape of the distribution, and that includes the one-sample case the signed-rank test belongs to. With data that are not normal it can be considerably more efficient than the t-test. If your data are genuinely interval and normal, the related t-test is slightly more powerful and should be used.

The Wilcoxon signed-rank test formula explained

The Wilcoxon signed-rank test has no complicated formula. T is the smaller of two rank totals: the sum of the ranks given to the positive differences, and the sum of the ranks given to the negative differences.

SymbolWhat it means
dThe difference for one pair: condition A score minus condition B score, always subtracted in the same order
NThe number of pairs left after removing any pair with d = 0
ΣR+The sum of the ranks belonging to positive differences
ΣR−The sum of the ranks belonging to negative differences
TThe smaller of ΣR+ and ΣR−
The symbols used in the Wilcoxon signed-rank test.

There is one built-in check, and it is worth using every time. The ranks 1 to N always add up to N(N + 1) ÷ 2, so the two rank totals must add up to that number. With 11 pairs, ΣR+ + ΣR− must equal 11 × 12 ÷ 2 = 66. If your totals do not, a rank has been missed or given twice.

Why the smaller total is the one that matters

The smaller total tells you how much evidence there is against the overall pattern. If every single difference were positive, the negative ranks would add up to zero, and T = 0 is the strongest possible result. If the conditions made no difference at all, the ranks would split roughly evenly and both totals would sit near N(N + 1) ÷ 4, which is 33 for 11 pairs. So the further T falls below that halfway point, the more lopsided the data and the less likely it is that chance alone produced it. That is why, as with the sign test and Mann-Whitney, a smaller number is the more impressive one.

Wilcoxon signed-rank test worked example: differences ranked by size, T = 6.5 with N = 11 against a critical value of 10
The whole test on one page. Notice that the two negative differences are small ones, ranked 1 and 5.5, which is why T ends up so low even though 11 pairs are involved.

How to calculate the Wilcoxon signed-rank test: a worked example

To calculate the Wilcoxon signed-rank test, find the difference for each pair, drop any zeros, rank the remaining differences by size ignoring the sign, add up the positive and negative ranks separately, and take the smaller total as T. Then compare T with the critical value for your N. The example below walks through every step, including a zero difference and tied ranks, because those are exactly the parts exam questions like to test.

A psychology student wants to know whether background music affects memory. Twelve participants each learn a list of 20 words in silence and a different, equally difficult list of 20 words with music playing. The order of the conditions is counterbalanced, so half the participants do silence first and half do music first. The number of words each person recalls is recorded. Because the same people do both conditions, this is a repeated measures design.

Step 1: State the hypotheses

The student has no strong reason to predict which way the effect will go, so the hypothesis is non-directional and the test is two-tailed.

  • Experimental hypothesis: There will be a difference in the number of words recalled from a 20-word list when learning in silence compared with learning with background music.
  • Null hypothesis: There will be no difference in the number of words recalled from a 20-word list when learning in silence compared with learning with background music.

The significance level is set at p ≤ 0.05, the conventional level in psychology. Our guide to statistical significance and p-values explains why that level is used and what it does and does not tell you.

Step 2: Find the difference for each pair

Take each participant’s music score away from their silence score. The order does not matter, but it must be the same for every pair, otherwise the signs become meaningless.

ParticipantSilence (A)Music (B)Difference (A − B)
11511+4
212120
3149+5
41011−1
51613+3
61311+2
7118+3
81710+7
9912−3
10148+6
111210+2
121512+3
Words recalled out of 20 by the same 12 participants in each condition.

A quick look already suggests a pattern. The median is 13.5 words in silence and 11 words with music, and only two participants did better with music. The test will tell us whether that pattern is strong enough to be unlikely by chance.

Step 3: Remove any zero differences

Participant 2 recalled 12 words in both conditions, so their difference is 0. A zero gives no evidence either way, so it is dropped before ranking and is not counted in N. That leaves N = 11, not 12. Forgetting this step is one of the most common ways to lose marks, because it changes both the ranks and the critical value.

Step 4: Rank the differences by size, ignoring the sign

Put the 11 remaining differences in order of size as if every one were positive, and give the smallest rank 1. Where two or more differences are the same size, they share the average of the ranks they would have taken. Then put each difference’s sign back on its rank.

Size of differenceParticipantsRank positionsRank givenSigned ranks
1411−1
26, 112 and 32.5+2.5, +2.5
35, 7, 9, 124, 5, 6 and 75.5+5.5, +5.5, −5.5, +5.5
4188+8
5399+9
6101010+10
781111+11
Ranking the 11 non-zero differences. The four differences of size 3 share the average of ranks 4 to 7, which is 5.5.

Notice that participant 9’s difference of −3 gets the same rank as the three differences of +3. Size decides the rank; the sign only decides which total the rank goes into. Notice too that the highest rank is 11, matching N. If your highest rank does not equal N, something has gone wrong.

Step 5: Add up the positive and negative ranks

Add up the ranks that carry a plus sign, then the ranks that carry a minus sign.

  • Positive ranks: 8 + 9 + 5.5 + 2.5 + 5.5 + 11 + 10 + 2.5 + 5.5 = 59.5
  • Negative ranks: 1 + 5.5 = 6.5
  • Check: 59.5 + 6.5 = 66, and 11 × 12 ÷ 2 = 66. The totals are correct.

Step 6: T is the smaller total

T is the smaller of 59.5 and 6.5, so T = 6.5, with N = 11.

Step 7: Find the critical value

Look up N = 11 in the critical values table, in the column for a two-tailed test at p ≤ 0.05. The critical value is 10.

Step 8: Compare and state the conclusion

For the Wilcoxon test, T must be equal to or less than the critical value to be significant. T = 6.5 is less than 10, so the result is significant at p ≤ 0.05. The null hypothesis is rejected and the experimental hypothesis is accepted: participants recalled significantly more words in silence than with background music.

It is worth taking one step further. At the stricter two-tailed level of p ≤ 0.01, the critical value for N = 11 is 5, and 6.5 is not equal to or less than 5, so the result is not significant at that level. The honest summary is therefore that the difference is significant at p ≤ 0.05 but not at p ≤ 0.01. Working out the exact probability puts it at roughly p = 0.014, which sits between the two.

Wilcoxon signed-rank test calculator and critical values table

This free calculator does every step above for you. Enter the scores for each condition in the same participant order, and it drops zeros, ranks the differences with averaged ties, works out T, and compares it with the exact critical value for your N. It also shows the exact probability, and switches to the normal approximation for N above 50. The critical values table underneath works without the calculator and covers N = 5 to 50.

Wilcoxon Signed-Rank Test Calculator

Type or paste each participant’s score in condition A and condition B, in the same order, separated by commas or spaces. Related designs only: repeated measures or matched pairs.

Critical values of T. The pink header row is the significance level for a one-tailed test and the dark row beneath it is the same column for a two-tailed test. N is the number of pairs left after dropping any with a difference of zero.

Critical values of T for the Wilcoxon signed-rank test, N = 5 to 50. T must be equal to or less than the value shown.
N0.050.0250.010.005
0.100.050.020.01
50–––
620––
7320–
85310
98531
1010853
11131075
12171397
132117129
1425211512
1530251915
1635292319
1741342723
1847403227
1953463732
2060524337
2167584942
2275655548
2383736254
2491816961
25100897668
26110988475
271191079283
2813011610191
29140126110100
30151137120109
31163147130118
32175159140128
33187170151138
34200182162148
35213195173159
36227208185171
37241221198182
38256235211194
39271249224207
40286264238220
41302279252233
42319294266247
43336310281261
44353327296276
45371343312291
46389361328307
47407378345322
48426396362339
49446415379355
50466434397373

A dash means no value of T can reach that level with so few pairs. Every value was calculated exactly from the distribution of T and is the largest T whose probability does not exceed the stated level.

Try the worked example in the calculator: it is already in the placeholder text. Entering 15, 12, 14, 10, 16, 13, 11, 17, 9, 14, 12, 15 for condition A and 11, 12, 9, 11, 13, 11, 8, 10, 12, 8, 10, 12 for condition B returns T = 6.5 with N = 11, exactly as calculated by hand.

A second worked example: when the result is not significant

A result that is not significant is calculated in exactly the same way, and seeing one helps you recognise what a weak pattern looks like. This example also shows the Wilcoxon test with rating-scale data, which produces a great many ties.

Ten students rate how anxious they feel on a scale of 1 to 10, then do a ten-minute breathing exercise and rate themselves again. A rating scale is ordinal: the gap between a 4 and a 5 is not guaranteed to feel the same as the gap between an 8 and a 9. The researcher predicts that anxiety will be lower after the exercise, so the hypothesis is directional and the test is one-tailed. The differences are calculated as before minus after, so a positive difference means anxiety went down, which is the predicted direction.

StudentBeforeAfterDifferenceSigned rank
165+1+3.5
2770dropped
356−1−3.5
486+2+8
545−1−3.5
664+2+8
776+1+3.5
856−1−3.5
964+2+8
1087+1+3.5
Anxiety ratings before and after a breathing exercise. Six differences of size 1 share ranks 1 to 6, averaging 3.5, and three differences of size 2 share ranks 7 to 9, averaging 8.

Student 2's zero is dropped, leaving N = 9. The positive ranks add up to 3.5 + 8 + 8 + 3.5 + 8 + 3.5 = 34.5, and the negative ranks add up to 3.5 + 3.5 + 3.5 = 10.5. The check works: 34.5 + 10.5 = 45, which is 9 × 10 ÷ 2. So T = 10.5.

The critical value for N = 9, one-tailed at p ≤ 0.05, is 8. T = 10.5 is greater than 8, so the result is not significant and the null hypothesis is retained. There is not enough evidence that the breathing exercise reduced anxiety.

Look at why. Six of the nine students who changed did feel calmer, which sounds promising, but every change was only one or two points and three students actually felt more anxious. With changes this small and this mixed, the ranks split too evenly to rule out chance. The exact one-tailed probability is about 0.08. This does not prove the exercise is useless. It means this study, with nine usable pairs and a coarse rating scale, could not detect an effect if there is one, which is a possible Type II error, and a larger sample would be the obvious next step.

How to read the Wilcoxon critical values table

To read the Wilcoxon critical values table, find the row for N, the number of pairs with a non-zero difference, then choose the column for your significance level and for whether your hypothesis is one-tailed or two-tailed. The number where they meet is the critical value, and your T must be equal to or less than it for the result to be significant.

NOne-tailed 0.05 / two-tailed 0.10One-tailed 0.025 / two-tailed 0.05One-tailed 0.01 / two-tailed 0.02One-tailed 0.005 / two-tailed 0.01
50---
620--
7320-
85310
98531
1010853
11131075
12171397
132117129
1425211512
1530251915
2060524337
25100897668
30151137120109
A quick-reference extract of the most used rows. The full table from N = 5 to 50 is in the calculator above.

Why T must be equal to or less than the critical value

T must be equal to or less than the critical value because a small T means very little evidence against the overall direction of the data. The critical value is the largest T that would happen by chance no more than 5% of the time, or whatever level you chose, if the two conditions really made no difference. Any T at or below it is rare enough to count as significant.

This is the opposite of the rule for chi-squared, Spearman's rho and the t-tests, where the calculated value must be equal to or greater than the critical value. The sign test, Wilcoxon and Mann-Whitney are the three tests where smaller is better. A useful memory aid: if the test letter is S, T or U, the value has to be low.

One-tailed or two-tailed?

Use the one-tailed columns when your hypothesis predicted a direction, and the two-tailed columns when it did not. Most printed tables have two header rows for this reason: each column has one significance level for a one-tailed test and double that level for a two-tailed test. So the column headed 0.025 for a one-tailed test is the same column as 0.05 for a two-tailed test.

With a one-tailed test there is an extra check. The data must actually go in the direction you predicted. If you predicted better recall in silence but most differences favoured music, a small T does not support your hypothesis, however significant it looks in the table. In that situation the directional hypothesis has to be rejected.

What the dashes in the table mean

A dash means that no value of T can be significant at that level with so few pairs. With N = 5, even the perfect result of T = 0 has a probability of 1 in 32, or about 0.031, one-tailed. That beats 0.05 but not 0.025, which is why a two-tailed test at p ≤ 0.05 is impossible with only five usable pairs. In practice, aim for at least eight to ten pairs that differ.

A note on the Edexcel table

The Pearson Edexcel specification prints a Wilcoxon table for N = 5 to 12, and every value in it matches the exact calculation except one. For N = 10 at one-tailed p ≤ 0.05, Edexcel prints 11, while the exact critical value is 10. The probability of T being 10 or less by chance is 0.042, which is inside the 0.05 limit, but the probability of T being 11 or less is 0.053, which is just outside it (Pearson Edexcel, 2026).

The difference only matters when T is exactly 11 with N = 10 on a one-tailed test. In an Edexcel exam, use the table you are given, because that is what the mark scheme follows. Everywhere else, 10 is the value to use. The calculator on this page flags this case if it comes up. The extract of the Wilcoxon table printed in AQA exam papers, for N = 19 to 22, matches the exact values on this page in every cell (AQA, n.d.).

How to deal with zero differences and tied ranks

Zero differences are removed before ranking and are not counted in N, and tied differences share the average of the ranks they would have occupied. Those two rules are what every A level specification expects, and they are the rules the tables are built for.

Zero differences. A pair with the same score in both conditions carries no information about which condition is better, so the standard method discards it. Pearson Edexcel spells this out in its instructions to candidates: do not rank any differences of 0, and do not count them when working out N (Pearson Edexcel, 2026). The drawback is that throwing data away makes significance a little harder to reach. Statisticians have proposed alternatives, the best known being Pratt's method, which ranks the zeros along with everything else and then drops their ranks (Pratt, 1959). At A level, simply dropping zeros is correct. If more than a handful of pairs are zeros, that is itself worth commenting on, because it may mean the measure is too coarse to pick up differences.

Tied ranks. When two or more differences are the same size, add up the rank positions they occupy and divide by how many there are. Two differences sharing positions 2 and 3 each get 2.5; four sharing positions 4 to 7 each get 5.5. The next difference then takes the next unused position, not the next whole number after the shared rank. This keeps the rank total equal to N(N + 1) ÷ 2, so the built-in check still works. Tied ranks can produce a T with a half in it, like the 6.5 in the worked example, and that is compared with the table in exactly the same way.

Heavy ties, like the anxiety ratings above, make the printed table only approximate, because it was worked out on the assumption that every difference is distinct. Statistics software applies a correction for ties when it calculates the probability. For A level hand calculation, the table is used as it stands.

Large samples: the z approximation

When N is too large for the printed table, T is converted into a z score and compared with the normal distribution. As N grows, the distribution of T becomes very close to a normal curve with a mean of N(N + 1) ÷ 4 and a standard deviation of the square root of N(N + 1)(2N + 1) ÷ 24.

The formula is z = (T − N(N + 1) ÷ 4) ÷ √(N(N + 1)(2N + 1) ÷ 24). For the worked example, the mean is 11 × 12 ÷ 4 = 33, the standard deviation is √(11 × 12 × 23 ÷ 24) = √126.5 = 11.25, and z = (6.5 − 33) ÷ 11.25 = −2.36. Ignoring the minus sign, 2.36 is bigger than 1.96, the two-tailed critical z at p ≤ 0.05, so the conclusion is the same as the table gave.

A level exam tables stop well before this is needed, and the approximation is rough for small samples, so use the exact table whenever N is covered by it. The z score matters mainly because statistics software reports it, and because it is used to calculate an effect size.

How to write up a Wilcoxon signed-rank test result

A Wilcoxon result is written up by naming the test, describing the direction of the difference with the medians of each condition, and then giving T, N, the significance level and whether the test was one-tailed or two-tailed. Medians are used rather than means because the test is based on ranks, and the median is the average that matches it.

For the worked example, a suitable write-up would be:

A Wilcoxon signed-rank test showed that participants recalled significantly more words when learning in silence (median = 13.5) than with background music (median = 11), T = 6.5, N = 11, p ≤ 0.05, two-tailed. The null hypothesis was rejected.

Three details earn or lose marks. Give N after zeros are removed, not the number of participants tested, and say how many pairs were dropped if any were. State the observed value and the critical value when the question asks you to explain your decision, for example "T = 6.5 is less than the critical value of 10". And put the conclusion in the words of the hypothesis, about words recalled and music, not only in terms of "the null hypothesis".

University reports in APA style usually add an exact p value and an effect size, and often report the z score from software as well. The next section shows how to get an effect size by hand.

Effect size for the Wilcoxon signed-rank test

The simplest effect size for the Wilcoxon signed-rank test is the matched-pairs rank-biserial correlation: the positive rank total minus the negative rank total, divided by the sum of all the ranks. Kerby (2014) showed that this is the same as the proportion of the rank evidence that favours the hypothesis minus the proportion that goes against it, which makes it easy to explain in words.

For the worked example, (59.5 − 6.5) ÷ 66 = 0.80. About 90% of the rank total points towards silence and about 10% towards music, and 0.90 − 0.10 = 0.80. The scale runs from −1 to +1 like any correlation coefficient, with 0 meaning no consistent difference and 1 meaning every pair changed the same way. This is a large effect.

The other common measure is r = z ÷ √N, using the z score from the normal approximation (Rosenthal, 1991). For the worked example that is 2.36 ÷ √11 = 0.71. Be careful with this one, because textbooks disagree about N: some divide by the number of pairs and others by the total number of observations, which is twice as many and gives a smaller r (Field, 2018). Always say which you used. The effect size matters because a significant result only tells you the difference is unlikely to be chance, not whether it is big enough to care about, a distinction our guide to standard error also explores.

The Wilcoxon signed-rank test sits between the sign test and the related t-test. All three compare two related conditions, and they differ in how much of each difference they use: the sign test uses only its direction, Wilcoxon uses its direction and its rank, and the t-test uses its actual size.

TestDesignMinimum level of dataWhat it uses from each pairStatisticSignificant when
Sign testRelatedNominalDirection onlySS ≤ critical value
Wilcoxon signed-rankRelatedOrdinalDirection and rank of sizeTT ≤ critical value
Related t-testRelatedInterval, roughly normalActual size of differencett ≥ critical value
Mann-Whitney UIndependent groupsOrdinalRank of each score in the pooled setUU ≤ critical value
How the Wilcoxon test compares with the tests it is most often confused with.

Wilcoxon vs the sign test

The Wilcoxon test and the sign test both handle related designs, but the sign test only counts pluses and minuses while Wilcoxon also takes account of how big each difference is. That usually makes Wilcoxon more powerful, meaning more likely to detect a real difference.

The worked example shows this clearly. Run through the sign test, the same data give 9 pluses and 2 minuses with N = 11, so S = 2. The two-tailed critical value of S for N = 11 at p ≤ 0.05 is 1, so the sign test says not significant, with an exact probability of about 0.065. Wilcoxon says significant, at about 0.014. The difference is that the sign test treats participant 4's tiny difference of −1 as just as important as participant 8's +7. Wilcoxon sees that both minuses were small and that the large differences all point towards silence.

It does not always work out that way, and the reason is instructive. In Darwin's maize data, discussed below, 13 of 15 differences were positive, but the two negative ones were large. The sign test gives a two-tailed probability of about 0.007 there, while Wilcoxon gives about 0.041. Wilcoxon is weighing the evidence rather than counting it, so big differences in the wrong direction cost more.

Wilcoxon vs the related t-test

The related t-test is the parametric version of the same comparison. It works on the mean of the differences and their spread, so it needs interval data and differences that are roughly normally distributed. When those conditions hold it is a little more powerful than Wilcoxon. When they do not, and particularly when there are outliers, Wilcoxon is the safer and often the more powerful choice (Hodges & Lehmann, 1956). Our full guide to t-tests in psychology covers the related t-test step by step.

Wilcoxon vs Mann-Whitney

Wilcoxon and Mann-Whitney are both rank-based tests of difference for ordinal data, and the only thing that separates them is the design. Wilcoxon is for related designs, where every score has a partner, so it ranks the differences within pairs. Mann-Whitney is for independent groups, where there are no pairs, so it ranks every score from both groups together. The names cause confusion because the Mann-Whitney U test is also called the Wilcoxon rank-sum test, since Wilcoxon's 1945 paper introduced both. "Signed-rank" always means the related version.

Wilcoxon vs Friedman

The Wilcoxon signed-rank test compares exactly two related conditions. If the same participants do three or more conditions, the Friedman test is the rank-based equivalent. Running several Wilcoxon tests instead of one Friedman test inflates the chance of a false positive, so with three conditions start with Friedman.

Checking the method against Wilcoxon's own examples and an AQA question

Wilcoxon's 1945 paper includes two paired examples with published answers, and an AQA exam question has a published mark scheme. All three make excellent practice, because you can check your working against a known result.

Wilcoxon's wheat experiment

Wilcoxon's first paired example came from a seed treatment experiment on wheat, with eight blocks each growing treatment A and treatment B (Wilcoxon, 1945). The differences between the stands of wheat were +58, +32, +30, +5, −7, +6, +11 and +10. Ranked by size, the only negative difference, −7, is the third smallest, so it gets rank 3.

That gives T = 3 with N = 8. The two-tailed critical value for N = 8 at p ≤ 0.05 is 3, so the result is significant, but only just. Wilcoxon reported that his table put the probability between 0.024 and 0.055. The exact two-tailed probability is 10 ÷ 256, or 0.039, which sits inside his range, and the calculator on this page returns the same figure.

Darwin's maize plants

Wilcoxon's second example used data Charles Darwin collected on pairs of maize plants, one cross-fertilised and one self-fertilised, grown in the same pot, which Ronald Fisher had already analysed with a t-test in The Design of Experiments (Darwin, 1876; Fisher, 1935). The 15 differences in height, measured in eighths of an inch, were 6, 8, 14, 16, 23, 24, 28, 29, 41, −48, 49, 56, 60, −67 and 75.

The two negative differences are the 10th and 14th largest, so T = 10 + 14 = 24, with N = 15. The two-tailed critical value for N = 15 at p ≤ 0.05 is 25, so the result is significant. Wilcoxon noted that Fisher's t-test had given a probability of 0.0497 for the same data (Wilcoxon, 1945, p. 80). The exact Wilcoxon probability is 0.041. Two very different methods, one using the full measurements and one using only their order, reach the same verdict, and that agreement is the whole argument for ranking methods.

An AQA exam question

In one AQA A level question, a psychologist predicted that memory test scores would be lower after a restricted diet than before it, analysed the data with the Wilcoxon signed ranks test, and obtained T = 53. Students were given an extract of the critical values table and asked whether the result was significant at the 5% level.

The hypothesis predicts a direction, so the test is one-tailed. The mark scheme gives the answer: the result is significant, because T = 53 is less than the critical value of 60 for N = 20 at p ≤ 0.05 one-tailed (AQA, n.d.). Stating "significant" earns one mark and the explanation earns two more, and a student who says "not significant" gets no marks at all, however good the explanation. The exact one-tailed probability for T = 53 with N = 20 is about 0.027.

Common Wilcoxon mistakes in exams

Most lost marks on the Wilcoxon test come from a small number of repeated mistakes, and every one of them is avoidable with a checklist.

  1. Ranking the raw scores instead of the differences. That is how Mann-Whitney works. Wilcoxon only ever ranks the differences.
  2. Keeping the signs when ranking. Rank by size alone, so −3 and +3 share a rank, then reattach the signs.
  3. Leaving zero differences in. Drop them, and reduce N to match.
  4. Using the number of participants as N. N is the number of pairs that differ, which is what the critical value depends on.
  5. Getting tied ranks wrong. Tied differences take the average of the positions they occupy, and the next difference takes the next unused position.
  6. Reversing the decision rule. T must be equal to or less than the critical value. Bigger is not better here.
  7. Using the wrong column. Check one-tailed against two-tailed and the significance level before reading the table.
  8. Forgetting the direction on a one-tailed test. A small T only supports a directional hypothesis if the data went the predicted way.
  9. Skipping the check. The positive and negative rank totals must add to N(N + 1) ÷ 2. It takes seconds and catches most arithmetic slips.

What each exam board requires

All three major A level Psychology specifications in England include the Wilcoxon test, but they expect different things of it, so it is worth knowing exactly what yours asks for. The AQA specification, for example, requires students to calculate only the sign test by hand, while Pearson Edexcel prints the full Wilcoxon procedure for candidates.

Exam boardWhat the specification saysWhat that means for you
AQA (7182)Section 4.2.3.3 lists Wilcoxon among the tests students must know when to use. Calculation is required only for the sign test.Know when to choose Wilcoxon and how to use its critical values table. Hand calculation is good practice rather than a requirement.
OCR (H567)Criteria for using, and understanding the use of, the Wilcoxon Signed Ranks test, alongside the other non-parametric tests.Know when to choose it, why, and how to interpret the result against a table.
Pearson Edexcel (9PS0)Mann-Whitney U and Wilcoxon are the named non-parametric tests of difference. The specification prints the Wilcoxon process and a critical values table.Be ready to follow the printed process to find T from raw data, and to use the test in the practical investigations where a related design fits.
Sources: AQA (2015), OCR (2026) and Pearson Edexcel (2026) specifications.

Whichever board you take, the choosing skill is examined every year, and it depends on recognising a related design. Pearson Edexcel's practical investigation for cognitive psychology asks students to use a Mann-Whitney U or Wilcoxon test "as appropriate", and suggests experiments such as a dual task study of working memory (Pearson Edexcel, 2026). Run as repeated measures, that study leads directly to Wilcoxon. For revision, practising the whole calculation once or twice by hand is the most reliable way to fix the steps in memory, and our guide to evidence-based revision methods explains why that kind of active practice works better than rereading notes.

Where the Wilcoxon test is used in psychology

The Wilcoxon signed-rank test is used wherever psychologists measure the same people twice, or compare carefully matched pairs, and cannot rely on the data being normal. That covers a great deal of real research.

  • Before and after studies. Ratings of mood, anxiety, confidence or pain before and after a therapy session, a training course or an intervention are ordinal and often skewed, which is exactly the job Wilcoxon was built for.
  • Memory experiments. Comparing recall under two conditions with the same participants, as in the silence and music example, is a standard way of testing predictions from models such as the multi-store model of memory, and practical projects of this kind usually have small samples.
  • Matched pairs and twin studies. Comparing each twin with their co-twin, or each participant with a matched control, produces natural pairs.
  • Clinical and single-group studies. Where only a small group is available, such as a case series of patients measured at two time points, a rank-based test avoids assumptions a handful of people cannot support.
  • Early years and education research. Children's scores on the same assessment at the start and end of a term, or observation ratings from two sessions, are often ordinal and come from small groups, so they suit Wilcoxon well.

Because it works on ranks, the Wilcoxon test also turns up whenever qualitative material has been coded into numbers, such as ratings given by judges to interview transcripts before and after an intervention.

Conclusion

The Wilcoxon signed-rank test answers one question: when the same people, or matched pairs, have been measured in two conditions, did one condition generally produce higher scores? Choose it for a difference, a related design and data that are at least ordinal. Find each difference, drop the zeros, rank the sizes while ignoring the sign, and total the positive and negative ranks. T is the smaller total, and it is significant when it is equal to or less than the critical value for N.

It is more powerful than the sign test because it weighs how big each difference is, and safer than the related t-test when data are ordinal or skewed. More than eighty years after a chemist devised it to save time on seed trials, it still gives the same answer as heavier methods from a fraction of the arithmetic. Use the calculator and table above to check your own working, and the worked examples to practise until the steps are automatic.

Frequently Asked Questions

What is the Wilcoxon signed-rank test used for?

It is used to test whether two related sets of scores differ, when the data are ordinal or not normally distributed. Typical uses are before-and-after measurements on the same people and comparisons between matched pairs. It is the non-parametric alternative to the related t-test.

What is the difference between the Wilcoxon signed-rank test and the sign test?

Both compare two related conditions. The sign test only records whether each pair went up or down, so it works with nominal data. Wilcoxon also ranks how large each change was, so it needs ordinal data but usually detects real differences more easily.

What is the difference between the Wilcoxon test and a paired t-test?

A paired t-test uses the actual sizes of the differences and assumes they are roughly normal and measured on an interval scale. Wilcoxon uses only their ranks, so it copes with ordinal data, skew and outliers. With normal interval data the t-test is slightly more powerful.

Is the Wilcoxon signed-rank test the same as the Mann-Whitney U test?

No. The signed-rank test is for related designs, where each score has a partner. Mann-Whitney is for two independent groups. Confusingly, Mann-Whitney is also called the Wilcoxon rank-sum test, but "signed-rank" always refers to the paired version.

What are the assumptions of the Wilcoxon signed-rank test?

The scores must come in related pairs, each pair must be independent of the others, and the differences must be at least ordinal so they can be ranked. Normality is not required. If you want to draw conclusions about the median difference specifically, the differences should also be roughly symmetrical.

How do you calculate the Wilcoxon signed-rank test by hand?

  1. Subtract each pair's scores in the same order.
  2. Discard zero differences and count the rest as N.
  3. Rank the differences by size, ignoring signs and averaging ties.
  4. Total the positive ranks and the negative ranks.
  5. Take the smaller total as T and compare it with the table.

What does the Wilcoxon test measure?

It measures how consistently one condition beats the other across pairs, weighted by how large each change is. T is the rank total on the minority side, so a low T means nearly all the evidence, especially the large differences, points the same way.

What is the difference between Wilcoxon signed-rank and Wilcoxon rank-sum?

Signed-rank is for paired or repeated measures data and ranks the differences within pairs. Rank-sum is for two separate groups and ranks all the scores together, and it is the same test as Mann-Whitney U. Frank Wilcoxon introduced both in one 1945 paper.

What sample size do you need for a Wilcoxon signed-rank test?

At least 6 non-zero pairs for a two-tailed test at p ≤ 0.05, because with 5 even a perfect result cannot reach significance. With 6 pairs only a perfect T of 0 is significant, so aim for 10 or more, where T can reach 8 and a few pairs may go the other way.

References

AQA. (2015). AS and A-level Psychology specification (7181, 7182). Assessment and Qualifications Alliance.

AQA. (n.d.). A-level Psychology: Inferential testing, questions and mark scheme. Assessment and Qualifications Alliance.

Conover, W. J. (1999). Practical nonparametric statistics (3rd ed.). John Wiley & Sons.

Darwin, C. (1876). The effects of cross and self fertilisation in the vegetable kingdom. John Murray.

Field, A. (2018). Discovering statistics using IBM SPSS Statistics (5th ed.). SAGE Publications.

Fisher, R. A. (1935). The design of experiments. Oliver and Boyd.

Hodges, J. L., & Lehmann, E. L. (1956). The efficiency of some nonparametric competitors of the t-test. The Annals of Mathematical Statistics, 27(2), 324-335.

Kerby, D. S. (2014). The simple difference formula: An approach to teaching nonparametric correlation. Comprehensive Psychology, 3, Article 11.IT.3.1.

OCR. (2026). A Level Psychology H567 specification (Version 1.5). Oxford Cambridge and RSA Examinations.

Pearson Edexcel. (2026). Specification: Pearson Edexcel Level 3 Advanced GCE in Psychology (9PS0) (Issue 4). Pearson Education Limited.

Pratt, J. W. (1959). Remarks on zeros and ties in the Wilcoxon signed rank procedures. Journal of the American Statistical Association, 54(287), 655-667.

Rosenthal, R. (1991). Meta-analytic procedures for social research (Rev. ed.). SAGE Publications.

Wilcoxon, F. (1945). Individual comparisons by ranking methods. Biometrics Bulletin, 1(6), 80-83.

Further Reading and Research

Recommended Articles

Suggested Books

  • Research Methods and Statistics in Psychology by Hugh Coolican
    • The standard UK text for psychology students moving from A level to degree study. It explains every A level test, including Wilcoxon, with worked examples and clear guidance on choosing between them.
  • Discovering Statistics Using IBM SPSS Statistics by Andy Field
    • An approachable and thorough statistics textbook. Its chapter on non-parametric tests shows how the Wilcoxon signed-rank test is run and reported in software, including effect sizes.
  • Practical Nonparametric Statistics by W. J. Conover
    • A rigorous reference for rank-based methods, setting out the assumptions, theory and exact tables behind the signed-rank test for readers who want to go beyond A level.

Recommended Websites

  • JSTOR: Wilcoxon (1945), Individual Comparisons by Ranking Methods
    • The original four-page paper that introduced the signed-rank test, with the wheat and maize examples and Wilcoxon's first probability table.
  • Pearson Edexcel Psychology qualification pages
    • The Edexcel specification prints the Wilcoxon Signed Ranks test process and critical values table in its appendix, alongside the Mann-Whitney formulae.
  • AQA A-level Psychology specification
    • Section 4.2.3.3 sets out which inferential tests students must know and what they must be able to do with each.

Kathy Brodie

Kathy Brodie is an Early Years Professional, Trainer and Author of multiple books on Early Years Education and Child Development. She is the founder of Early Years TV and the Early Years Summit.

Kathy's Author Profile
Kathy Brodie

To cite this article please use:

Early Years TV Wilcoxon Signed-Rank Test: Formula, Worked Example and Table. Available at: https://www.earlyyears.tv/wilcoxon-signed-rank-test-formula-worked-example-table/ (Accessed: 2 October 2026).