T-Tests in Psychology: A Complete Guide

Roughly one in four p-values published across major psychology journals comes from a t-test — making it the single most-used statistical test for comparing two groups in psychological research today.
Key Takeaways:
- What is a t-test used for? It tells you whether the difference between two group averages is likely to be real, or just chance — used constantly in psychology to compare treatment vs. control groups or before-and-after scores.
- Independent or paired — which do I need? If the same people were measured twice, use a paired-samples t-test; if they’re separate groups of people, use an independent-samples t-test.
- How do I report a t-test in APA format? State the means and standard deviations for each group, then report t(df) = , p = , alongside an effect size such as Cohen’s d.
What Is a T-Test?
A t-test is a statistical method used to work out whether the difference between the average scores of two groups is likely to be a real, meaningful difference, or whether it could simply have happened by chance. In psychology, where researchers are constantly comparing groups — a treatment group against a control group, scores before an intervention against scores after it, or one population against another — the t-test is one of the most frequently reached-for tools in the entire statistical toolkit.
At its heart, a t-test does something quite intuitive. It looks at the size of the difference between two group means (the “signal”) and compares this against how much the scores within each group naturally vary (the “noise”). If the difference between the groups is large relative to the natural spread of scores within them, the t-test suggests the difference is unlikely to be down to chance alone. If the difference is small compared to the natural variation, the t-test suggests we cannot be confident the groups are really different at all.
The result of a t-test is a “t-statistic” (often just called “t”), which is then converted into a p-value — a number that tells us how likely we would be to see a difference this large, or larger, if there were actually no real difference between the groups. Researchers typically treat a p-value below 0.05 as evidence that the difference is statistically significant, meaning it is unlikely to have arisen by chance alone.
Because it is relatively simple to calculate, easy to interpret, and works well with the small-to-moderate sample sizes common in psychology, the t-test remains a cornerstone of quantitative research methods, taught early in every psychology degree and used throughout professional research careers.
Why Psychology Relies on the T-Test
Psychology is fundamentally a science of comparison. Does a new therapy reduce anxiety more than no treatment at all? Do memory scores improve after a good night’s sleep compared to after sleep deprivation? Do two age groups differ in reaction time? Almost every one of these questions boils down to comparing the average score of one group against another — exactly the situation the t-test was built for.
This explains why the t-test appears so often in the psychological literature. One analysis of p-values published across eight major psychology journals between 1985 and 2013 found that out of more than 258,000 p-values reported, roughly a quarter were testing a t-statistic — making it one of the single most commonly reported statistics in the entire field, alongside the F-statistic used in ANOVA (Nuijten et al., 2016, as cited in Gronau et al., 2017). Another review found an average of over three t-tests reported per psychology journal article (Wetzels et al., 2011, as cited in Gronau et al., 2017) — underlining just how embedded this test is in everyday psychological research practice.
Its popularity isn’t accidental. The t-test is mathematically well suited to situations with two groups and relatively modest sample sizes, which describes a huge proportion of psychological studies, from small-scale lab experiments to classroom-based research projects. It also produces results that translate neatly into the standardised scores that underpin so much of psychological measurement — the same logic that underlies the Z-score, another standardisation tool psychologists rely on constantly, is at work behind the scenes of every t-test calculation.

The Three Types of T-Test
Although people often talk about “the” t-test as if it were a single test, there are actually three distinct versions, each suited to a different kind of research question and data structure.
The one-sample t-test compares the mean of a single group against a known or hypothesised value. For example, a researcher might want to know whether the average score on a new wellbeing questionnaire, taken by a sample of teachers, differs significantly from a previously established population average.
The independent-samples t-test (also called the unpaired or between-subjects t-test) compares the means of two separate, unrelated groups. This is the test to reach for when the participants in Group A are entirely different people from the participants in Group B — for instance, comparing exam anxiety scores between students who received a mindfulness intervention and students who did not.
The paired-samples t-test (also called the dependent, related, or within-subjects t-test) compares two sets of scores that come from the same participants, or from participants who are closely matched. The classic example is measuring the same group of people twice — once before an intervention and once after — and testing whether their scores changed significantly.
Of the three, the independent-samples and paired-samples versions are by far the most commonly used in real research, and knowing which one applies to your data is one of the most important — and most commonly confused — decisions a student or researcher has to make.
Each type maps onto a distinct kind of research design that psychology students will encounter throughout their studies. The one-sample t-test tends to appear when researchers have access to well-established population norms — for example, national averages on standardised cognitive or personality measures — and want to check whether a specific sample (perhaps a clinical group, or a particular demographic) differs from that established benchmark. The independent-samples t-test is the natural choice for between-subjects experimental designs, where different participants are randomly allocated to different conditions, making it a staple of laboratory-based psychology experiments. The paired-samples t-test, meanwhile, is the natural fit for within-subjects designs, where the same people are tracked across time or across two different conditions — common in clinical and educational psychology, where measuring change within the same individuals is often more meaningful than comparing different people to one another.
Understanding which design a study actually used is therefore the first and most important step before any statistics are run at all — the choice of t-test isn’t really a statistical decision so much as a direct consequence of how the study was designed in the first place.
Independent-Samples vs Paired-Samples T-Test: The Decision Tree
Choosing between these two tests comes down to a single question: are the two sets of scores you’re comparing linked to each other, or are they completely separate?
Ask yourself the following, in order:
- Are the same participants contributing scores to both groups? If yes — for example, the same people were tested before and after an intervention — you are dealing with related data, and a paired-samples t-test is appropriate.
- If different participants make up each group, are they matched or paired in some deliberate way (such as identical twins, or participants matched on age and IQ before being split into two conditions)? If so, this also calls for a paired-samples t-test, because the pairing creates a meaningful link between specific scores in each group.
- If neither of the above applies — the participants in each group are simply different individuals with no deliberate matching — then the groups are independent of one another, and an independent-samples t-test is the correct choice.
| Feature | Independent-Samples T-Test | Paired-Samples T-Test |
|---|---|---|
| Participants | Two separate, unrelated groups | Same participants (or matched pairs) measured twice |
| Typical design | Between-subjects (e.g., treatment vs. control) | Within-subjects (e.g., before vs. after) |
| Example question | Do Group A and Group B differ in anxiety scores? | Did anxiety scores change from before to after treatment? |
| What’s compared | Mean difference between two independent groups | Mean of the differences within pairs |
| Statistical power | Generally lower, for the same total sample size | Generally higher, because individual variability is controlled for |
This last point about statistical power is worth remembering: because a paired design measures the same people twice, it automatically accounts for each person’s own baseline characteristics, which reduces “noise” in the comparison and makes it easier to detect a genuine effect. This is one reason researchers often prefer a repeated-measures design when it’s practical and ethical to use one.
That said, a paired design isn’t always possible or appropriate. Testing the same participants twice can introduce its own complications, such as practice effects (participants performing better simply because they’ve encountered the task before) or fatigue effects (performance declining due to tiredness or reduced motivation on a second attempt). In situations where taking a measurement twice would meaningfully change the participant — for example, testing knowledge before and after they’ve already seen the answers — an independent-samples design comparing two separate groups may actually be the more valid choice, even though it typically requires a larger overall sample to achieve the same statistical power. Choosing between the two designs, in other words, is rarely just a statistical decision; it also depends on practical and methodological considerations specific to the research question at hand.
Worked Example — Independent-Samples T-Test
Imagine a researcher wants to know whether a short mindfulness course reduces exam-related anxiety in university students, compared to students who receive no such course. Ten students are randomly assigned to the mindfulness group, and a separate ten students act as a control group who receive no intervention. All twenty students then complete a standard anxiety questionnaire, scored out of 50 (higher scores indicate greater anxiety).
The mindfulness group’s scores are: 22, 25, 20, 24, 21, 23, 19, 26, 22, 20 (mean = 22.2)
The control group’s scores are: 30, 28, 33, 29, 27, 31, 32, 26, 30, 29 (mean = 29.5)
At first glance, the mindfulness group’s average anxiety score (22.2) is noticeably lower than the control group’s (29.5) — a difference of 7.3 points. But is this difference large enough, relative to the natural spread of scores within each group, to be confident it reflects a genuine effect of the mindfulness course rather than random variation between two groups of different people?
To understand where the t-statistic actually comes from, it helps to walk through the logic behind the calculation, even though in practice almost all researchers use statistical software rather than working through it by hand. The independent-samples t-test formula essentially divides the difference between the two group means by a measure of how much the scores vary within each group (known as the standard error of the difference):
t = (Mean₁ – Mean₂) / Standard Error of the Difference
In this example, the mindfulness group has a standard deviation of approximately 2.20, and the control group has a standard deviation of approximately 2.12 — reassuringly similar, which supports the assumption of equal variance discussed later in this article. Combining these into the standard error and dividing it into the 7.3-point difference between the means produces a t-statistic of approximately t = 6.02, with 18 degrees of freedom (calculated as the total number of participants across both groups minus two, since two separate group means are being estimated from the data). Looking this value up against the t-distribution shows that a t-statistic this large corresponds to a p-value well below 0.001.
Because this p-value is far smaller than the conventional 0.05 threshold, the researcher can conclude that the difference in anxiety scores between the mindfulness group and the control group is statistically significant — it is extremely unlikely that a difference this large would occur simply by chance if the mindfulness course had no real effect. In plain terms: the students who took the mindfulness course reported meaningfully lower exam anxiety than those who didn’t, and this gap is unlikely to be a fluke of who happened to be assigned to each group.
It’s worth pausing on what this result does, and doesn’t, tell us. A significant t-test confirms that a difference between the groups is unlikely to be due to chance — but it doesn’t, by itself, tell us how large or practically important that difference is. That’s a separate question, answered by effect size, which is covered in the reporting section below.
Worked Example — Paired-Samples T-Test
Now imagine a different design for the same underlying question. Instead of comparing two separate groups, the researcher measures the same ten students’ anxiety scores twice: once before the mindfulness course, and again after they complete it.
Before the course: 30, 28, 33, 29, 27, 31, 32, 26, 30, 29 (mean = 29.5)
After the course: 22, 25, 20, 24, 21, 23, 19, 26, 22, 20 (mean = 22.2)
Notice this uses the same numbers as the two groups above, but reframed as a single group measured twice. A paired-samples t-test doesn’t just compare the two overall means — it looks at the difference score for each individual student (for example, student one went from 30 to 22, a drop of 8 points), and then tests whether the average of these individual differences is significantly different from zero.
The average difference across all ten students in this example is 7.3 points, and because the same people are being compared to themselves, the natural person-to-person variability that exists between two separate groups is removed from the equation entirely. The paired-samples t-test formula reflects this directly — rather than comparing two separate group means, it divides the average of the individual difference scores by the standard error of those differences:
t = Mean of the Differences / Standard Error of the Differences
Because every student’s “before” score is being compared only to their own “after” score, any stable individual traits — a naturally anxious personality, or a naturally calm one — cancel out of the calculation entirely. This typically produces a larger t-statistic for the same underlying difference in means, compared to the equivalent independent-samples test — in this case, the paired-samples t-test produces t = 13.86 with 9 degrees of freedom (calculated as the number of pairs minus one, rather than the number of participants minus two), and again a p-value well below 0.001.
This comparison illustrates why the choice between independent and paired designs matters so much: the paired design, by controlling for each individual’s own baseline, is statistically more powerful and can detect real effects more easily — one of the practical reasons repeated-measures designs are so popular in psychology whenever they are feasible.
A common mistake worth flagging here: researchers sometimes run an independent-samples t-test on what is actually paired data (for instance, treating “before” and “after” scores as if they came from two unrelated groups). Doing so ignores the fact that the same individuals contributed both scores, discards useful information about person-to-person variability, and typically produces a less accurate — and less statistically powerful — result than the correctly matched paired-samples test.
Assumptions of the T-Test
Like all parametric statistical tests, the t-test relies on the data meeting certain conditions in order for its results to be trustworthy. Three assumptions matter most:
Normality. The scores within each group should be roughly normally distributed — that is, they should form something close to the familiar bell-shaped curve, with most scores clustering around the average and progressively fewer scores further away from it in either direction. With larger sample sizes, the t-test becomes fairly robust to modest departures from normality, but with small samples, a badly skewed distribution can distort the results.
Homogeneity of variance (for the independent-samples t-test specifically). This assumption requires that the two groups being compared have roughly similar amounts of spread, or variability, in their scores. If one group’s scores are tightly clustered while the other group’s scores are widely spread out, the standard independent-samples t-test can become unreliable. Where this assumption is clearly violated, researchers typically switch to a variation called Welch’s t-test, which adjusts the calculation to account for unequal variances.
Independence of observations. For the independent-samples t-test, each participant’s score must not influence or be linked to any other participant’s score — this is why random assignment to groups is so important in experimental designs. For the paired-samples t-test, by contrast, the pairs themselves are expected to be related (that is the whole point of the design), but the different pairs should still be independent of one another.
In practice, researchers check these assumptions before running the t-test itself, often using a histogram or normality test to inspect the shape of the data, and a test such as Levene’s test to check for equal variances between groups. A useful discussion of what happens to test conclusions when these assumptions are met — or not — can be found in the University of Southern Queensland’s open statistics resource for research students (Statistics for Research Students, n.d.).
It’s also worth understanding what tends to happen when these assumptions are violated in real research. With independence violations — for example, if participants in an “independent” group discussed the study with each other and influenced one another’s responses — the resulting p-value can become unreliable in either direction, because the statistical model assumes each data point carries genuinely separate information. With normality violations in small samples, the t-test can become overly sensitive to a small number of extreme or unusual scores, which can distort the mean and inflate or deflate the apparent difference between groups. This is one reason why researchers are often encouraged to visually inspect their data — using simple tools like histograms or box plots — rather than relying purely on the test result itself.
Fortunately, the t-test is often described as reasonably “robust” to minor or moderate assumption violations, particularly once sample sizes grow beyond around 20–30 participants per group, thanks to a statistical principle known as the central limit theorem, which describes how sample means tend to become more normally distributed as sample size increases, even when the underlying data itself is not perfectly normal. This doesn’t mean assumptions can be ignored altogether, but it does explain why the t-test remains usable across a very wide range of real psychological datasets, rather than being restricted only to perfectly textbook-shaped data.
Reading and Reporting T-Test Results in APA Format
Once a t-test has been run, the result needs to be reported clearly and consistently so that other researchers can understand exactly what was tested and how strong the evidence was. The American Psychological Association’s (APA) referencing and reporting style sets out a standard format for this.
A typical APA-style report of an independent-samples t-test result looks like this:
“An independent-samples t-test was conducted to compare anxiety scores between the mindfulness group and the control group. There was a significant difference in scores for the mindfulness group (M = 22.2, SD = 2.20) and the control group (M = 29.5, SD = 2.12); t(18) = 6.02, p < .001.”
Breaking this down: t(18) tells the reader the t-statistic (6.02) and the degrees of freedom (18) on which it is based. p < .001 tells the reader the probability of observing a difference this large, or larger, purely by chance, if there were truly no difference between the groups — in this case, a very small probability, supporting the conclusion of a genuine effect.
A paired-samples result follows a very similar format, but is described slightly differently to reflect the repeated-measures design:
“A paired-samples t-test was conducted to compare anxiety scores before and after the mindfulness course. Anxiety scores were significantly lower after the course (M = 22.2, SD = 2.20) than before it (M = 29.5, SD = 2.12); t(9) = 13.86, p < .001.”
It’s good academic practice to report an effect size alongside the p-value, since a p-value alone does not indicate how large or meaningful the difference actually is — only how confident we can be that a difference exists at all. This distinction matters more than it might first appear: with a large enough sample size, even a tiny, practically unimportant difference between two groups can produce a statistically significant p-value, while a genuinely large and meaningful difference can fail to reach significance if the sample is too small.
Cohen’s d is the most commonly reported effect size for t-tests. It expresses the difference between the two group means in standardised units, calculated by dividing the mean difference by the pooled standard deviation of the two groups — conceptually very similar to the standardisation logic behind a Z-score. Conventionally, a Cohen’s d of around 0.2 is considered a small effect, around 0.5 a medium effect, and 0.8 or above a large effect (Lakens, 2013). In the mindfulness example above, the 7.3-point difference relative to the groups’ standard deviations produces a Cohen’s d well above 2 — an unusually large effect size, reflecting how cleanly separated the two groups’ scores are in this illustrative example.
A complete APA-style report therefore typically includes the effect size alongside the t-statistic and p-value, for example: “t(18) = 6.02, p < .001, d = 2.69.” Including this figure allows other researchers, and anyone reading a summary of the findings, to judge not just whether an effect exists, but how substantial it is — information that matters a great deal when decisions about clinical practice, educational policy, or further research funding are being made. The statistical significance and effect sizes guide covers how to calculate and interpret Cohen’s d, along with Type I and Type II errors, in more depth.
When Not to Use a T-Test
The t-test is designed specifically for comparing the means of two groups (or two sets of related scores). This is its most important limitation: it simply isn’t built for situations involving three or more groups.
If a researcher wanted to compare exam anxiety across three different intervention groups — for example, mindfulness, exercise, and a control group — running multiple separate t-tests (mindfulness vs. control, exercise vs. control, mindfulness vs. exercise) would inflate the overall risk of a false positive result, because each additional test carries its own small chance of error. In this situation, an ANOVA (analysis of variance) is the appropriate test, since it is designed to compare three or more group means in a single, statistically sound analysis.
Similarly, the t-test is intended for continuous, interval- or ratio-scale data — such as scores on a questionnaire, reaction times, or physiological measurements. If the data of interest is categorical instead (for example, comparing the proportion of participants in two groups who passed versus failed a task), a chi-square test is the appropriate tool rather than a t-test.
Finally, where the assumptions discussed earlier are seriously violated — particularly with small samples that show clear non-normality — a non-parametric alternative such as the Mann-Whitney U test (for independent samples) or the Wilcoxon signed-rank test (for paired samples) may give more trustworthy results than a standard t-test.
There is also a subtler limitation worth flagging: repeatedly running separate t-tests across many different comparisons within the same dataset — even when each comparison technically only involves two groups — carries the same inflated false-positive risk described above. A researcher comparing five different outcome measures between two groups, for instance, running five separate t-tests rather than a single combined analysis, substantially increases the chance that at least one of those five results will appear significant purely by chance. Where multiple comparisons are unavoidable, statistical corrections (such as the Bonferroni correction) or a different overall test are generally the more defensible approach.
Recognising these boundaries is just as important as knowing how to run the test itself — using the right tool for the number of groups and type of data being compared is fundamental to producing research findings that hold up to scrutiny.
Conclusion
The t-test remains one of the most widely used tools in psychological research precisely because it answers a question researchers ask constantly: is the difference between two groups real, or could it simply be down to chance? Choosing correctly between an independent-samples and a paired-samples design, checking the test’s underlying assumptions, and reporting results clearly in APA format are the practical skills that turn a simple calculation into trustworthy evidence. Used appropriately — and recognising when a different test, such as ANOVA or chi-square, is needed instead — the t-test continues to underpin a huge proportion of the psychological research published every year.
Frequently Asked Questions
What is the t-test used for in psychology?
A t-test determines whether the difference between the average scores of two groups is likely to reflect a genuine effect, or simply chance variation. Psychologists use it to compare treatment and control groups, scores before and after an intervention, or two distinct populations on any continuous measure. It is one of the most widely reported statistics in psychological research, underpinning conclusions about whether an intervention, condition, or characteristic produces a measurable difference in outcomes.
What does the t-test value tell you?
The t-value (or t-statistic) shows how large the difference between two group means is relative to the variability within the groups. A larger t-value means the difference is large compared to the natural spread of scores, making it less likely to have occurred by chance. Combined with the degrees of freedom, the t-value is converted into a p-value, which researchers use to judge whether a result counts as statistically significant.
When do you use a t-test vs ANOVA?
Use a t-test when comparing the means of exactly two groups. Use ANOVA (analysis of variance) when comparing three or more groups. Running several separate t-tests instead of one ANOVA increases the risk of a false positive result, since each extra comparison carries its own chance of error. ANOVA avoids this by testing all groups together in a single model, making it the more reliable choice whenever more than two groups are involved.
When do you use a t-test vs a Z-test?
Both compare group means, but a t-test is generally used when the population standard deviation is unknown and sample sizes are small to moderate — which describes most psychological research. A Z-test is used when the population standard deviation is known, or when sample sizes are very large (typically over 30), since the t-distribution converges toward the normal distribution as sample size grows. In practice, psychologists use t-tests far more often than Z-tests.
What are the different types of t-test?
There are three main types. The one-sample t-test compares a single group’s mean against a known value. The independent-samples t-test compares the means of two separate, unrelated groups. The paired-samples t-test compares two sets of related scores, typically the same participants measured twice. Choosing the right type depends on how the data was collected — specifically, whether the two sets of scores come from the same people or from separate groups.
References
- Gronau, Q. F., Ly, A., & Wagenmakers, E.-J. (2017). Informed Bayesian t-tests. arXiv:1704.02479.
- Lakens, D. (2013). Calculating and reporting effect sizes to facilitate cumulative science: A practical primer for t-tests and ANOVAs. Frontiers in Psychology, 4, 863.
- Statistics for Research Students. (n.d.). Section 3.4: Paired t-test assumptions, interpretation, and write up. University of Southern Queensland Open Access Textbooks.
Further Reading and Research
Recommended Articles
- Delacre, M., Lakens, D., & Leys, C. (2017). Why psychologists should by default use Welch’s t-test instead of Student’s t-test. International Review of Social Psychology, 30(1), 92–101.
- Ruxton, G. D. (2006). The unequal variance t-test is an underused alternative to Student’s t-test and the Mann–Whitney U test. Behavioral Ecology, 17(4), 688–690.
- Lakens, D. (2013). Calculating and reporting effect sizes to facilitate cumulative science: A practical primer for t-tests and ANOVAs. Frontiers in Psychology, 4, 863.
Suggested Books
- Field, A. (2017). Discovering statistics using IBM SPSS statistics (5th ed.). SAGE Publications.
- A comprehensive statistics textbook covering t-tests, ANOVA, and regression, with worked examples using SPSS throughout.
- Howell, D. C. (2012). Statistical methods for psychology (8th ed.). Cengage Learning.
- An in-depth treatment of statistical methods used in psychological research, including detailed coverage of t-tests and their assumptions.
- Coolican, H. (2019). Research methods and statistics in psychology (7th ed.). Routledge.
- A widely used student text covering research design and statistical analysis, including when and how to apply t-tests.
Recommended Websites
- Statistics How To — Offers clear explanations of t-tests, worked examples, and free calculators for independent and paired designs.
- UCLA Institute for Digital Research and Education — Provides detailed guides on running and interpreting t-tests in statistical software such as SPSS and R.
- StatTrek — Offers free tutorials and step-by-step guidance on t-test calculations and hypothesis testing.
To cite this article please use:
Early Years TV T-Tests in Psychology: A Complete Guide. Available at: https://www.earlyyears.tv/t-tests-in-psychology-a-complete-guide/ (Accessed: 7 August 2026).

