Experimental Design in Psychology: 3 Types With Examples

Comparison of independent groups, repeated measures and matched pairs designs: strengths, weaknesses and controls

In a between-subjects experiment, one group of people judged the number 9 to be significantly larger than another group judged 221 (Birnbaum, 1999). Nothing was wrong with the arithmetic. The experimental design produced the absurd result.

Key Takeaways

  • Experimental design is how participants are allocated to the conditions of an experiment. There are three: independent groups (different people in each condition), repeated measures (the same people in every condition) and matched pairs (different but paired people).
  • Each design trades one problem for another. Independent groups suffers from participant variables, repeated measures suffers from order effects and demand characteristics, and matched pairs is slow and can never match people perfectly.
  • The design decides the statistical test. Repeated measures and matched pairs are related designs; independent groups is an unrelated design, and the tests differ for each.

Every experiment in psychology compares at least two conditions. Before a single participant arrives, the researcher has to decide who goes into which condition. Should the same people do everything, or should each condition have its own people? That decision is the experimental design, and it shapes almost everything that follows: how many participants are needed, which problems can creep in, how the results are analysed and how far they can be trusted.

For A level students, experimental design is one of the most heavily examined ideas in research methods. It turns up in short “identify the design” questions, in “explain one strength and one limitation” questions, and inside almost every question about choosing the right statistical test, because the test depends on the design as much as it depends on the level of measurement.

This guide explains all three designs in plain language, with real studies for each one, the strengths and weaknesses examiners expect, how to deal with order effects through counterbalancing, how many participants each design needs, and a practice quiz to test yourself at the end.

What Is Experimental Design in Psychology?

Experimental design in psychology is the way participants are allocated to the different conditions of the independent variable. The independent variable (IV) is the thing the researcher changes, such as whether music is playing or not. The conditions are its different levels, such as “music” and “silence”. The design answers one question: does each person take part in one condition, or all of them?

The AQA A level specification lists exactly three: “repeated measures, independent groups, matched pairs” (AQA, 2015). Edexcel lists the same three names, and OCR uses slightly different labels for the same ideas. University textbooks and research papers use a fourth set of terms again, which is why the same design can appear under several names.

AQA and EdexcelOCRUniversity and research papersWho takes part in each condition
Independent groupsIndependent measuresBetween-subjects, between-groupsDifferent people
Repeated measuresRepeated measuresWithin-subjectsThe same people
Matched pairsMatched participantsMatched-subjects, matched groupsDifferent people, paired on key characteristics

Experimental design should not be confused with the type of experiment. Lab, field, natural and quasi-experiments describe where a study happens and how much control the researcher has over the IV. Experimental design describes who is in each condition. A lab experiment can use any of the three designs, and so can a field experiment.

It also only applies to experiments. A correlation compares two co-variables measured on the same people, so there are no conditions to allocate anyone to. Methods such as a case study or an interview do not have an experimental design at all.

Why Experimental Design Matters

Experimental design matters because an experiment only works if the conditions differ in one way: the IV. Anything else that differs between the conditions is a possible alternative explanation for the result. Psychologists call these extraneous variables, and when one of them actually changes along with the IV it becomes a confounding variable.

The two biggest sources of confounding in any experiment are the people and the order. If different people are in each condition, their differences in ability, personality or mood might explain the result. If the same people do every condition, what happened in the first condition might change how they perform in the second. Each experimental design is a different answer to that dilemma.

Comparison of independent groups, repeated measures and matched pairs designs: strengths, weaknesses and controls
No design wins on every row. Each one solves the main problem of another and pays for it somewhere else, which is exactly what examiners want you to explain.

The Three Experimental Designs Compared

The three experimental designs differ in who takes part in each condition, and that single difference sets their strengths and weaknesses. The table below is the quickest way to revise them side by side.

Independent groupsRepeated measuresMatched pairs
Who is in each conditionDifferent people, randomly allocatedThe same people do every conditionDifferent people, paired on a relevant variable
Participant variablesA problemRemovedReduced
Order effectsNoneA problemNone
Demand characteristicsLess likelyMore likelyLess likely
Participants neededThe mostThe fewestAs many as independent groups, plus a pre-test
Main fixRandom allocationCounterbalancingCareful matching on a pre-test
Related or unrelated dataUnrelatedRelatedRelated

What Is an Independent Groups Design?

An independent groups design uses different participants in each condition of the experiment. Each person experiences only one level of the IV, and the results of one group are compared with the results of the other. OCR calls this an independent measures design, and research papers usually call it a between-subjects design.

Independent Groups Example: Loftus and Palmer (1974)

The second experiment in Loftus and Palmer’s famous car crash study is a textbook independent groups design. One hundred and fifty students watched a short film of a multiple car accident. Fifty were then asked how fast the cars were going when they “smashed into” each other, fifty were asked the same question with “hit”, and fifty were not asked about speed at all (Loftus and Palmer, 1974).

A week later, everyone was asked whether they had seen any broken glass. There was none in the film. Sixteen of the 50 people in the “smashed” group said yes, compared with 7 in the “hit” group and 6 in the control group (Loftus and Palmer, 1974). Because every student was in only one group, each person’s answer could only have been shaped by one version of the question.

That is the point of the design. Had the same students been asked both the “smashed” and the “hit” question, the first question would have contaminated the second, and the study would have made no sense.

Random Allocation

Random allocation means every participant has an equal chance of ending up in each condition. It is the main control for an independent groups design, and the AQA specification names it alongside counterbalancing as a control technique students must know (AQA, 2015).

Random allocation does not make the groups identical. What it does is make any differences between them a matter of chance rather than bias. Without it, a researcher might, without meaning to, put the keener volunteers into the condition they expect to do well. A simple way to randomly allocate is:

  1. Give every participant a number.
  2. Put the numbers in a hat, or use a random number generator.
  3. Assign the first half drawn to condition A and the second half to condition B.

Random allocation is not the same as random sampling. Random sampling is about who gets into the study from the target population. Random allocation is about which condition they go into once they are in. A study can use an opportunity sample and still randomly allocate that sample to conditions.

Strengths of an Independent Groups Design

  • No order effects. Each participant only does one condition, so practice, boredom and fatigue from an earlier condition cannot affect their performance.
  • Fewer demand characteristics. Seeing only one condition makes it harder for a participant to work out what the study is testing and change their behaviour to fit it.
  • The same materials can be used in every condition. In a memory experiment, both groups can learn the same word list, so any difference cannot be caused by one list being easier.
  • It is quick to run. Each person takes part once, and there is no need for a second session or a pre-test.

Weaknesses of an Independent Groups Design

  • Participant variables. The people in one group may differ from the people in the other in ways that affect the dependent variable (DV), such as memory ability, age or motivation. If the “music” group happens to contain better learners, the result reflects the people rather than the music. This is the biggest weakness, and random allocation reduces it but never removes it.
  • More participants are needed. To get the same amount of data, the researcher needs a separate set of people for every condition. The section on sample size below shows how large that cost can be.
  • Differences can be harder to detect. Because natural variation between people adds noise to the comparison, a real effect of the IV has to be larger, or the sample bigger, before it shows up as statistically significant.

Between-subjects designs also bring a subtler problem, and it is the one behind the hook at the top of this article. When each group sees only one condition, each group judges it against its own frame of reference. Birnbaum (1999) showed that when different groups rated how large a single number was, 9 was rated significantly larger than 221. The explanation he offered is that 9 brings to mind small numbers, among which it seems fairly large, while 221 brings to mind three-digit numbers, among which it seems small. The same people judging both numbers would never have made that mistake.

That specific result has since been challenged. Leong and colleagues (2019) found that the 9 > 221 effect reversed when people rated on a 1 to 1000 scale or a slider rather than the original 1 to 10 scale. But they also found new effects of the same kind, including people rating 9 as larger than 009, and concluded that Birnbaum’s warning about comparing between-subjects ratings still stands.

What Is a Repeated Measures Design?

A repeated measures design uses the same participants in every condition of the experiment. Each person experiences every level of the IV, and their performance in one condition is compared with their own performance in the other. Research papers call this a within-subjects design.

A pre-test and post-test study is a common kind of repeated measures design. Measuring the anxiety of the same group of clients before and after a course of therapy compares each person with themselves.

Repeated Measures Example: Godden and Baddeley (1975)

Godden and Baddeley’s divers study is the classic repeated measures experiment in A level psychology, and one of the best-known pieces of evidence for context-dependent forgetting. Members of a university diving club listened to lists of 36 unrelated words, either sitting at the water’s edge or about 20 feet underwater, then recalled them either in the same place or in the other one (Godden and Baddeley, 1975).

Every diver took part in all four combinations: learn on land and recall on land, learn on land and recall underwater, learn underwater and recall on land, and learn underwater and recall underwater. Each condition was run in a separate diving session, roughly 24 hours apart. The mean scores out of 36 were:

LearnedRecalled on landRecalled underwater
On land13.58.6
Underwater8.411.4
Mean words recalled out of 36 (Godden and Baddeley, 1975, Table 1).

Recall was better when the place of recall matched the place of learning. Because each diver provided a score in every cell, differences in memory ability between divers could not explain the pattern. A diver with a poor memory scored low everywhere, and a diver with a good memory scored high everywhere, but each person’s own pattern still favoured the matching environment.

The researchers also dealt with the obvious problems of doing the same task four times. They used four different word lists, one per condition, so no one learned the same words twice, and they varied which list went with which condition and in what order, using a design called a Graeco-Latin square. That is counterbalancing, which is explained in full below.

A warning worth knowing: a 2021 replication with sixteen divers, again doing all four combinations, did not find better recall in the matching context (Murre, 2021). One study not replicating does not make the original wrong, but it is a good evaluation point in an exam.

Strengths of a Repeated Measures Design

  • Participant variables are removed. The same people are compared with themselves, so differences in intelligence, age, memory or personality between conditions are no longer possible. This is the design’s single biggest advantage.
  • Fewer participants are needed. Every participant provides data for every condition, so the researcher recruits one group instead of two or more.
  • It is statistically more sensitive. Removing the variation between people makes a real effect of the IV easier to detect. Greenwald (1976) described this statistical efficiency as the general principle favouring within-subjects designs.

Weaknesses of a Repeated Measures Design

  • Order effects. Doing one condition can change how people do the next, through practice, boredom or tiredness. This is the design’s biggest weakness and gets its own section below.
  • Demand characteristics. After experiencing every condition, participants are more likely to guess the aim of the study and change their behaviour, either to help the researcher or to spoil the results.
  • Different materials are often needed. A participant cannot learn the same word list twice without the second attempt being easier, so a second list is needed, and the two lists must be equally difficult. If they are not, the list becomes a confounding variable.
  • Some effects cannot be undone. Once someone has been taught to read, learned a strategy or been told the purpose of a deception, they cannot return to the untreated condition. For these IVs repeated measures is simply impossible.
  • Drop-out. When conditions are run on different days, some participants do not come back, and their data from the first session may be lost.

Order Effects and Counterbalancing

Order effects are changes in participants’ performance caused by the order in which they do the conditions, rather than by the IV. They only affect repeated measures designs, because only in that design does anyone do more than one condition. Edexcel lists counterbalancing and order effects as a named topic of their own (Pearson Edexcel, 2015), and AQA names counterbalancing as a control (AQA, 2015).

Types of Order Effect

  • Practice effect. Performance improves in later conditions because participants have become familiar with the task, the equipment or the kind of question.
  • Fatigue effect. Performance gets worse in later conditions because participants are tired.
  • Boredom effect. Performance gets worse because participants have lost interest by the second or third time round.
  • Carry-over effect. Something specific from one condition, such as a drug still in the bloodstream or a strategy learned in condition A, affects performance in condition B.

Order effects matter because they are confounded with the IV. If everyone does the “silence” condition first and the “music” condition second, and scores are higher with music, there is no way to tell whether the music helped or whether people had simply had some practice.

What Is Counterbalancing?

Counterbalancing is a way of arranging the order of conditions so that order effects are spread equally across them. It does not stop order effects from happening. It makes sure they affect every condition to the same extent, so they cannot favour one condition over another. There are two common forms.

AB/BA counterbalancing. The sample is split in half. Half the participants do condition A then condition B, and the other half do B then A. Any practice effect now helps A for one half and B for the other, so it balances out across the whole sample.

ABBA counterbalancing. Every participant does the conditions in the order A, B, B, A. Each condition appears once early and once late, so each person’s average score for A and for B contains the same amount of practice. It works when the task can be repeated several times, and assumes practice builds up steadily from trial to trial.

MethodWhat happensBest forMain limitation
AB/BAHalf the sample does A then B, half does B then ATasks that can only be done once or twiceOnly balances order across the group, not within each person
ABBAEvery participant does A, B, B, AShort tasks that can be repeatedTakes longer, so fatigue and boredom increase
Latin squareEach condition appears in each position once across groupsThree or more conditionsNeeds as many groups as conditions
Randomised orderEach participant gets a random orderMany short trials, as in reaction-time tasksOrder effects even out only across large numbers

Counterbalancing With Three or More Conditions

With three or more conditions, running every possible order quickly gets out of hand: three conditions have six possible orders, and four have twenty-four. A Latin square is the usual shortcut. It uses as many orders as there are conditions, arranged so each condition appears in each position exactly once. With three conditions A, B and C, three groups would do ABC, BCA and CAB.

Godden and Baddeley (1975) used a more elaborate version, a Graeco-Latin square, to balance two things at once: the order of the four conditions and which of the four word lists went with each condition. Their divers were split at random into four groups of four, and each group met the conditions and the lists in a different order.

When Counterbalancing Does Not Work

Counterbalancing only works if the effect of doing A before B is the same size as the effect of doing B before A. When it is not, psychologists call it asymmetrical transfer, and Poulton and Freeman (1966) showed it can produce misleading results even in balanced designs. Imagine a condition that teaches participants a useful memory strategy. People who do it first carry the strategy into the second condition, but people who do it second have nothing to carry back. Averaging the two orders does not cancel that out.

When a strong, one-way carry-over is likely, the better answer is usually not to use repeated measures at all.

Other Ways to Reduce Order Effects

  • Leave a time gap between conditions so fatigue and boredom wear off. Godden and Baddeley ran one condition per day.
  • Use different but equivalent materials in each condition, such as matched word lists, so participants cannot simply remember the previous set.
  • Include practice trials before the real conditions start, so most of the practice effect happens before any data are recorded.
  • Vary the order between groups, as Loftus and Palmer (1974) did in their first experiment: all participants watched seven different crash films, and each group of participants saw the films in a different order.

What Is a Matched Pairs Design?

A matched pairs design uses different participants in each condition, but first pairs them up on a characteristic that could affect the DV. One member of each pair goes into condition A and the other into condition B, usually at random. The pairs are then compared, so each comparison is between two people who are as alike as possible on the variable that matters. OCR calls this a matched participants design.

Matched pairs is an attempt to get the best of both other designs: no order effects, as in independent groups, and far fewer participant variables, closer to repeated measures.

Matched Pairs Example: Bandura, Ross and Ross (1961)

Bandura’s Bobo doll study, the foundation of social learning theory’s account of aggression, matched children before allocating them. The participants were 36 boys and 36 girls from the Stanford University Nursery School, aged 37 to 69 months, and the original paper describes the matching step by step (Bandura et al., 1961).

Before the experiment, the experimenter and a nursery teacher who knew the children well rated each child’s everyday aggression on four five-point scales: physical aggression, verbal aggression, aggression towards objects and aggressive inhibition. The two raters’ scores correlated at .89, so the ratings were reliable. On the basis of these scores the children were arranged in threes and assigned at random to one of three groups: an aggressive model, a non-aggressive model, or no model (Bandura et al., 1961).

This shows that “matched pairs” is not limited to pairs. With three conditions, matching produces matched triplets. The logic is identical: children who were already more aggressive were spread evenly across the conditions, so pre-existing aggression could not explain why the aggressive-model group behaved more aggressively afterwards.

How to Match Participants

  1. Decide which participant variable is most likely to affect the DV. For a memory experiment it might be memory span; for an aggression study, existing aggression.
  2. Measure every participant on that variable with a pre-test, such as a short memory test or a rating scale.
  3. Rank the participants on their pre-test scores.
  4. Pair the top two, then the next two, and so on down the list.
  5. Randomly allocate one member of each pair to each condition, for example by tossing a coin.

Identical twins are the closest thing to a perfect match, because they share their genes and usually their upbringing. Twin studies that put one twin in each condition are a natural form of matched pairs design.

Strengths of a Matched Pairs Design

  • No order effects. Each person only does one condition.
  • Participant variables are reduced. The variable most likely to affect the results is controlled by the matching.
  • Fewer demand characteristics than repeated measures. Participants see only one condition, so they have less to go on in guessing the aim.
  • The same materials can be used in both conditions, as in independent groups.

Weaknesses of a Matched Pairs Design

  • Matching is never perfect. People can be matched on one or two variables, but they still differ on everything else. Even identical twins differ in experience.
  • It is time-consuming and expensive. Every participant needs a pre-test before the real experiment can begin, and some pre-tests, such as Bandura’s teacher ratings, need people who know the participants well.
  • Participants are lost. If one member of a pair drops out, the other’s data can no longer be compared, and anyone without a good match cannot be used.
  • The pre-test may reveal the aim. Being tested on memory before a memory experiment can hint at what the study is about.
  • It needs as many participants as independent groups, because each person still only contributes to one condition.

How Many Participants Does Each Design Need?

A repeated measures design usually needs far fewer participants than an independent groups design to detect the same effect. Revision guides often say it needs half as many, because one group does the work of two. The real saving is usually larger than that, and it depends on how consistent people are from one condition to the next.

Cohen (1992) showed that an independent groups experiment needs 64 participants in each group, 128 in total, to have an 80% chance of detecting a medium-sized difference at the 0.05 level. The table below compares that with the number a repeated measures design needs for the same effect, calculated with the standard power formula for a related t-test. The key figure is the correlation between people’s scores in the two conditions: the more consistent people are, the more a repeated measures design gains.

DesignMedium effect (d = 0.5)Large effect (d = 0.8)
Independent groups (total, both groups)12852
Repeated measures, scores correlate 0.34620
Repeated measures, scores correlate 0.53415
Repeated measures, scores correlate 0.72110
Participants needed for 80% power, two-tailed test at p = 0.05. Independent groups figures from Cohen (1992); repeated measures figures calculated for this article.

With a medium effect and a moderate correlation of 0.5, repeated measures needs 34 people against 128, roughly a quarter. This is exactly what Greenwald (1976) meant by statistical efficiency: removing the differences between people from the comparison leaves less noise for the effect to overcome. It is also why the standard deviation of the differences, rather than of the raw scores, is what matters in a related design.

A matched pairs design sits between the two. It needs as many people as independent groups to recruit, but if the matching variable really does predict the DV, the pairs behave like correlated scores and the analysis gains some of the same sensitivity.

Which Experimental Design Is Best?

No experimental design is best in general. The best design is the one whose weakness does the least damage to the particular study. A useful way to decide is to ask three questions in order.

  1. Can the same person do every condition without the first one changing them? If not, because the IV involves learning, deception or a lasting change, repeated measures is ruled out. Use independent groups or matched pairs.
  2. Are participant variables likely to swamp the effect? If people differ a lot on the thing being measured, such as reaction time, memory or reading ability, and the effect is expected to be small, repeated measures is the strongest choice, provided counterbalancing can handle the order effects.
  3. Is there one participant variable that clearly matters and can be measured in advance? If so, and repeated measures is not possible, matched pairs controls that variable. If not, independent groups with random allocation is the simplest honest option.
SituationUsually the best designWhy
Testing a new teaching methodIndependent groups or matched pairsLearning cannot be undone
Comparing reaction times with and without a distractionRepeated measuresReaction times vary widely between people
Only a small number of participants availableRepeated measuresEvery person contributes to every condition
The study involves deception about its aimIndependent groupsA second condition would reveal the aim
A known variable, such as IQ, strongly affects the DVMatched pairsControls that variable without order effects
Comparing men and women, or two age groupsIndependent groupsNo one can be in both groups

The last row is worth noticing. When the IV is a characteristic people already have, such as age, gender or handedness, the researcher cannot allocate anyone to a condition at all, so the design is always independent groups. Studies like this are quasi-experiments, because the groups were not created by random allocation.

Greenwald (1976) made a further point that is easy to miss. Between-subjects designs are not free of context effects; each group simply has the context of its single condition. If the real world involves people meeting several conditions, such as a shopper comparing several prices, a repeated measures design may actually be the more realistic, more ecologically valid choice.

How to Identify the Design of a Study

To identify the design of a study, ask one question: did any single participant provide data in more than one condition of the IV? If yes, it is repeated measures. If no, check whether the participants were paired on a characteristic before allocation. If they were, it is matched pairs; if not, it is independent groups.

Do not be misled by the word “groups”. Studies that counterbalance often split participants into groups, but those groups exist to vary the order, not to separate the conditions. Godden and Baddeley (1975) is the clearest case. Their divers were “split at random into four groups of four”, which leads some revision sources to describe the study as independent groups. The same paper states that each group experienced every condition in a different order, and analysed the results with a Wilcoxon matched-pairs signed-ranks test, a test for related data. It is repeated measures.

Some studies mix designs. Loftus and Palmer’s first experiment used different people for each verb, an independent groups comparison, but every participant watched all seven films. When a study has two IVs, identify the design separately for each.

Experimental Design and Choosing a Statistical Test

Experimental design decides whether the data are related or unrelated, and that decides which statistical test to use. AQA states that the choice of test depends on “level of measurement and experimental design” (AQA, 2015).

  • Related data come from repeated measures and matched pairs designs. Each score in condition A has a partner in condition B, either the same person or their matched pair, so the analysis works on the differences within each pair.
  • Unrelated data come from independent groups designs. There is no natural partner for any score, so the analysis compares the groups as wholes.

A matched pairs design is always treated as related, even though different people are in each condition. This is one of the most common mistakes in exam answers.

Level of measurementRelated design (repeated measures, matched pairs)Unrelated design (independent groups)
NominalSign testChi-squared
OrdinalWilcoxonMann-Whitney
IntervalRelated t-testUnrelated t-test

Each of these tests has its own worked guide: the sign test, the Wilcoxon signed-rank test, the Mann-Whitney U test, the chi-squared test and the related and unrelated t-tests. The two classic studies above show the link in action: Godden and Baddeley’s repeated measures design was analysed with a Wilcoxon test, and Loftus and Palmer’s independent groups data on broken glass were analysed with a chi-squared test, significant beyond the .025 level (Loftus and Palmer, 1974).

Spot the Design: Practice Quiz

The quickest way to get confident with experimental designs is to classify real examples. Read each scenario, choose the design, and check the explanation. Several are real studies from the A level course.

Spot the Design

Ten scenarios. Pick the experimental design for each one and read why.

1. Thirty students learn a word list in silence and are tested. The following week the same students learn a different list with music playing and are tested again.

The same students provide data in both conditions. The second list avoids a practice effect from relearning the same words, but the design has no counterbalancing, so any improvement could still be practice.

2. Forty volunteers are randomly split into two groups. One group drinks caffeinated coffee and the other decaffeinated, then everyone does a reaction time test.

Different people are in each condition and they were randomly allocated, with no pairing beforehand.

3. Participants take an IQ test. The two highest scorers are split between the two conditions, then the next two, and so on down the list.

Participants are paired on a pre-test score before one of each pair goes into each condition. That is the matching procedure.

4. Loftus and Palmer (1974), experiment 2: after a crash film, 50 students are asked a question using “smashed”, 50 using “hit”, and 50 are not asked about speed.

Each student heard only one version of the question, so each condition had its own participants.

5. Godden and Baddeley (1975): divers are split into four groups of four, and every diver learns and recalls word lists in all four land and underwater combinations, in different orders.

Every diver did every condition. The four groups only exist to counterbalance the order, which is why this study is so often misclassified.

6. Bandura, Ross and Ross (1961): children are rated for everyday aggression, arranged in threes with similar ratings, and one of each three is assigned at random to each of three conditions.

Children were matched on aggression before allocation. With three conditions the matched sets are triplets rather than pairs, but the design is the same.

7. A researcher compares how many words left-handed and right-handed people can recall from the same list.

No one can be in both groups, so the design has to be independent groups. Because handedness cannot be randomly allocated, this is also a quasi-experiment.

8. Children rate two sets of book illustrations. Half see set A first and half see set B first.

Every child rates both sets. Splitting them into two halves is AB/BA counterbalancing, not a separate group for each condition.

9. In each of 15 pairs of identical twins, one twin revises with flashcards and the other with mind maps, then both sit the same test.

Different people in each condition, but each twin is paired with their sibling. Identical twins are the closest thing to a perfect match.

10. Twelve clients complete an anxiety questionnaire before and after eight weeks of therapy.

A pre-test and post-test on the same people is repeated measures. Order cannot be counterbalanced here, because “after therapy” can never come first.

Score: 0 out of 0 answered.

How to Answer Experimental Design Exam Questions

Experimental design exam questions reward answers that are tied to the study in the question rather than recited from a revision list. The same point about participant variables earns more when it names the actual variable that could differ between the groups.

“Identify the Experimental Design”

Name the design and give the reason from the scenario: “Repeated measures, because the same 20 students were tested both with and without music.” The reason shows the examiner you have not guessed.

“Explain One Strength or Limitation”

Use a three-step structure: state the point, apply it to this study, and say why it matters for the results. For example: “A limitation of the independent groups design is participant variables. The students in the music group may have had better memories than those in the silence group. This means any difference in recall might be due to memory ability rather than the music, which reduces internal validity.”

“How Could the Researcher Deal With This Problem?”

Match the fix to the design. Participant variables in independent groups: random allocation, or switch to matched pairs. Order effects in repeated measures: counterbalancing, with a description of how, such as “half the participants do the music condition first and half do the silence condition first.” Imperfect matching in matched pairs: match on the variable most likely to affect the DV, using a pre-test.

“Design a Study”

Name the design, justify it for this particular IV, explain how participants will be allocated or how order will be counterbalanced, and make sure the design fits the statistical test you choose. A plan that uses repeated measures and then an unrelated test contradicts itself.

Conclusion

Experimental design is the decision about who takes part in each condition of an experiment, and there are three answers. Independent groups uses different people and risks participant variables. Repeated measures uses the same people and risks order effects and demand characteristics. Matched pairs uses different people paired on a key variable, and costs time and effort in matching that is never perfect.

The controls follow from the problems: random allocation for independent groups, counterbalancing for repeated measures, and careful pre-testing for matched pairs. The design also sets the statistics, because repeated measures and matched pairs produce related data while independent groups produces unrelated data. Classic studies show all of this at work, from Loftus and Palmer’s separate groups to Godden and Baddeley’s counterbalanced divers and Bandura’s matched nursery children, and the Birnbaum result is a reminder that even the choice between designs can change what a study appears to find.

Frequently Asked Questions

What are the 3 experimental designs in psychology?

The three experimental designs in psychology are independent groups, repeated measures and matched pairs. Independent groups puts different participants in each condition. Repeated measures puts the same participants through every condition. Matched pairs uses different participants who have been paired on a relevant characteristic, with one of each pair in each condition. OCR calls them independent measures, repeated measures and matched participants.

What is the difference between repeated measures and independent groups?

Repeated measuresIndependent groups
ParticipantsSame people in all conditionsDifferent people in each condition
Main problemOrder effectsParticipant variables
Main controlCounterbalancingRandom allocation
DataRelatedUnrelated

Is matched pairs related or unrelated?

A matched pairs design produces related data. Although different people take part in each condition, every score has a partner: the score of the person they were matched with. Statistical tests therefore treat matched pairs in the same way as repeated measures, using the Wilcoxon test, the sign test or the related t-test rather than the tests for independent groups.

Why would a researcher use a matched pairs design?

A researcher uses a matched pairs design when repeated measures is impossible or would cause order effects, but one participant variable is likely to affect the results. Matching on that variable controls it without anyone doing two conditions. It suits studies where the first condition would permanently change participants, such as teaching a new skill.

What is counterbalancing in psychology?

Counterbalancing in psychology is a control for order effects in repeated measures designs. The order of conditions is varied so that practice, fatigue and boredom affect each condition equally. In AB/BA counterbalancing half the participants do condition A first and half do condition B first. Counterbalancing spreads order effects evenly rather than removing them.

What are order effects in psychology?

Order effects in psychology are changes in how participants perform that are caused by the order of the conditions rather than the independent variable. Practice can make later conditions easier, while tiredness and boredom can make them harder. Order effects only arise in repeated measures designs, where each person takes part in more than one condition.

What is the difference between repeated measures and matched pairs?

In repeated measures the same person provides both scores; in matched pairs two similar people provide one score each. Repeated measures controls every participant variable but suffers order effects. Matched pairs avoids order effects but only controls the variables used for matching. Both produce related data and use the same statistical tests.

Is repeated measures the same as within-subjects?

Yes. Repeated measures and within-subjects are two names for the same design, in which every participant experiences every condition. A level specifications use “repeated measures”, while university textbooks and journal articles more often say “within-subjects”. In the same way, independent groups is also called between-subjects.

References

  • AQA. (2015). AS and A-level Psychology specification (7181, 7182). AQA.
  • Bandura, A., Ross, D., and Ross, S. A. (1961). Transmission of aggression through imitation of aggressive models. Journal of Abnormal and Social Psychology, 63(3), 575-582.
  • Birnbaum, M. H. (1999). How to show that 9 > 221: Collect judgments in a between-subjects design. Psychological Methods, 4(3), 243-249.
  • Cohen, J. (1992). A power primer. Psychological Bulletin, 112(1), 155-159.
  • Godden, D. R., and Baddeley, A. D. (1975). Context-dependent memory in two natural environments: On land and underwater. British Journal of Psychology, 66(3), 325-331.
  • Greenwald, A. G. (1976). Within-subjects designs: To use or not to use? Psychological Bulletin, 83(2), 314-320.
  • Leong, L. M., McKenzie, C. R. M., Sher, S., and Müller-Trede, J. (2019). Illusory inconsistencies in judgment: Stimulus-evoked reference sets and between-subjects designs. Psychonomic Bulletin and Review, 26(2), 647-653.
  • Loftus, E. F., and Palmer, J. C. (1974). Reconstruction of automobile destruction: An example of the interaction between language and memory. Journal of Verbal Learning and Verbal Behavior, 13(5), 585-589.
  • Murre, J. M. J. (2021). The Godden and Baddeley (1975) experiment on context-dependent memory on land and underwater: A replication. Royal Society Open Science, 8(11), 200724.
  • OCR. (2015). A Level Psychology H567 specification. OCR.
  • Pearson Edexcel. (2015). Pearson Edexcel Level 3 Advanced GCE in Psychology (9PS0) specification. Pearson.
  • Poulton, E. C., and Freeman, P. R. (1966). Unwanted asymmetrical transfer effects with balanced experimental designs. Psychological Bulletin, 66(1), 1-8.

Further Reading and Research

Recommended Articles

Suggested Books

  • Coolican, H. (2019). Research Methods and Statistics in Psychology (7th ed.). Routledge.
    • The standard A level and undergraduate reference, with detailed chapters on experimental designs, counterbalancing and the threats each design is open to.
  • Field, A., and Hole, G. (2003). How to Design and Report Experiments. Sage.
    • A practical guide to planning an experiment from choosing a design through to writing it up, aimed at psychology students.
  • Shadish, W. R., Cook, T. D., and Campbell, D. T. (2002). Experimental and Quasi-Experimental Designs for Generalized Causal Inference. Houghton Mifflin.
    • The classic advanced text on what experiments can and cannot show, including the threats to validity behind every design choice.

Recommended Websites

  • AQA A-level Psychology (7182)
    • The specification, past papers and mark schemes, including the research methods content on experimental designs and controls.
  • Royal Society Open Science
    • Free full text of Murre’s 2021 replication of Godden and Baddeley, a clear example of a counterbalanced repeated measures design written up in full.
  • Classics in the History of Psychology (York University)
    • Free full texts of classic papers, including Bandura, Ross and Ross (1961), where the matching procedure can be read in the original.

Kathy Brodie

Kathy Brodie is an Early Years Professional, Trainer and Author of multiple books on Early Years Education and Child Development. She is the founder of Early Years TV and the Early Years Summit.

Kathy’s Author Profile
Kathy Brodie

To cite this article please use:

Early Years TV Experimental Design in Psychology: 3 Types With Examples. Available at: https://www.earlyyears.tv/experimental-design-psychology-repeated-independent-matched/ (Accessed: 2 October 2026).