Observational Techniques in Psychology: Types and Design

Chart of the six observational techniques in psychology in three pairs, with example studies and observational design steps

When eight sane people secretly had themselves admitted to psychiatric hospitals, staff never spotted them, but 35 of 118 patients counted on the wards did. That study is covert participant observation, one of psychology’s six observational techniques.

Key Takeaways

  • What are observational techniques? Ways of watching and recording what people actually do. The AQA specification names three pairs: naturalistic or controlled, covert or overt, and participant or non-participant. Every observation is one of each pair at the same time.
  • How is an observation designed? The researcher turns the behaviour into clear behavioural categories, then chooses event sampling (record every time it happens) or time sampling (record at fixed intervals).
  • How do you know it is reliable? Two or more observers watch the same behaviour with the same categories and compare their records. A strong positive correlation, or a high percentage agreement, shows inter-observer reliability.

Most research methods in psychology ask people what they think or do. Observation is different: it watches what they do. That matters, because what people say and what people do are often not the same thing. A parent may describe a calm bedtime routine, while a researcher watching the same household sees something quite different. Observation gets closer to real behaviour, which is why it sits behind some of psychology’s most famous studies, from Ainsworth’s Strange Situation to Bandura’s Bobo doll research.

The AQA A level specification lists “observational techniques” and “observational design” as two separate topics (AQA, 2015). The techniques are the types of observation. The design is how the observation is organised: what counts as the behaviour, when it is recorded, and how the researcher checks that the records can be trusted. Both link closely to other research methods topics, especially the types of experiment, because observation is often the way the dependent variable in an experiment is measured, and reliability, because two observers can easily see the same event in different ways.

This guide explains each of the six types of observation with a real example and its strengths and weaknesses, then walks through observational design step by step: behavioural categories, event and time sampling, and inter-observer reliability. It finishes with the common problems, a practice quiz for classifying observations and advice on answering exam questions.

What Are Observational Techniques in Psychology?

Observational techniques in psychology are methods of collecting data by watching and recording behaviour, rather than asking people about it. The AQA specification groups them as three pairs: “naturalistic and controlled observation; covert and overt observation; participant and non-participant observation” (AQA, 2015). Together these are often called the six types of observation.

Each pair answers a different question about the study:

  • Naturalistic or controlled? Where does the observation happen, and has the researcher arranged the situation? A naturalistic observation watches behaviour where it normally happens. A controlled observation takes place in a situation the researcher has set up, often a laboratory.
  • Covert or overt? Do the people being watched know? In a covert observation they do not. In an overt observation they know they are being observed.
  • Participant or non-participant? Is the researcher part of the group? A participant observer joins in with the people being studied. A non-participant observer watches from outside the group.

The key point, and the one students most often miss, is that these are not six separate methods. Every observation takes one option from each pair at the same time. Rosenhan’s (1973) study of psychiatric hospitals was naturalistic, covert and participant. Ainsworth’s Strange Situation was controlled, overt and non-participant. An exam question that asks you to “identify the type of observation” usually expects you to deal with more than one pair.

Observation as a Research Method and as a Technique

Observation can be a research method in its own right or a technique used inside another method. When a researcher simply watches children in a playground and describes the bullying they see, the observation is the whole study. Nothing is manipulated, so it cannot show cause and effect, only what happens and how often.

Often, though, observation is how an experiment measures its dependent variable. In Bandura, Ross and Ross’s (1961) Bobo doll study, the independent variable was the type of adult model the children saw. The dependent variable, the children’s aggression, was measured by observers watching through a one-way mirror and recording behaviour in set categories. The study was a laboratory experiment that used a controlled, non-participant observation to collect its data. This is why the same study can be described correctly as an experiment and as an observation: one describes the design, the other describes how the data were gathered.

Observation is also one of the main sources of data in a case study, where a researcher may watch one person or one family over a long period alongside interviews and records.

Chart of the six observational techniques in psychology in three pairs, with example studies and observational design steps
Read the chart one column at a time: every observation takes one option from each pair, so a single study always has three labels.

The Six Types of Observation Compared

The six types of observation trade realism against control and ethics against honesty of behaviour. Naturalistic, covert and participant observations tend to capture behaviour as it really is, while controlled, overt and non-participant observations tend to be easier to repeat, easier to record and easier to justify ethically.

TypeWhat it meansMain strengthMain weaknessExample
NaturalisticBehaviour watched where it normally happensHigh ecological validityLittle control over extraneous variablesCraig and Pepler (1998), playground bullying
ControlledSituation set up and standardised by the researcherEasy to replicateBehaviour may be artificialAinsworth and Bell (1970), Strange Situation
CovertParticipants do not know they are observedNo participant reactivityNo informed consentRosenhan (1973)
OvertParticipants know they are observedMore ethicalPeople may change their behaviourParten (1932), nursery play
ParticipantResearcher joins the groupInsight from the insideResearcher may lose objectivityFestinger et al. (1956)
Non-participantResearcher watches from outsideMore objectiveMay miss what behaviour meansBandura et al. (1961)

Keep in mind that the strengths and weaknesses are tendencies, not rules. A covert observation removes reactivity only if the covering story works, and a participant observer can stay objective with careful training and structured records. Good exam answers explain why a strength applies to the particular study in the question rather than listing it in general.

Naturalistic vs Controlled Observation

The difference between naturalistic and controlled observation is who decides the situation. In a naturalistic observation the researcher records behaviour in its usual setting and does not interfere. In a controlled observation the researcher arranges and standardises the situation, so every participant faces the same conditions.

What Is a Naturalistic Observation?

A naturalistic observation is one in which behaviour is studied in the environment where it would normally occur, such as a home, a classroom, a playground or a workplace, without the researcher changing anything. The researcher’s job is to record what happens, not to make it happen.

A clear example is Craig and Pepler’s (1998) study of bullying in school playgrounds. The researchers observed children during ordinary playtimes and identified bullying episodes with 90% agreement between raters. Bullying turned out to happen roughly once every seven minutes and to last only 38 seconds on average. Adults intervened in just 4% of episodes and peers in 11%, even though peers were involved in some way in 85% of them (Craig and Pepler, 1998). No questionnaire could have produced those figures, because the children involved would not have been able to report them accurately, and teachers did not see most of the episodes.

A follow-up compared playgrounds with classrooms using the same naturalistic approach and found bullying episodes happened more often in the playground, 4.5 per hour, than in the classroom, 2.4 per hour. The type of bullying also matched the setting: direct bullying was more common outside and indirect bullying more common in class (Craig, Pepler and Atlas, 2000). That finding only exists because the behaviour was watched in both of its real settings.

Strengths of naturalistic observation:

  • High ecological validity. Behaviour is recorded in its real setting, so findings are more likely to generalise to everyday life. See the guide to validity in psychology for why this matters.
  • It can study behaviour that cannot be recreated. Bullying, crowd behaviour or play between young children would be unethical or impossible to stage in a laboratory.
  • It can generate hypotheses. Patterns spotted in the real world, such as where bullying happens, can later be tested more tightly.

Weaknesses of naturalistic observation:

  • Little control. Many extraneous variables, such as the weather, which children are present or what happened earlier that day, can affect behaviour, so it is hard to explain why something happened.
  • Hard to replicate. The exact situation will never happen again, which makes it difficult to check the findings by repeating the study.
  • The behaviour may not happen. The researcher must wait for it, which can make naturalistic observations slow and costly.

What Is a Controlled Observation?

A controlled observation is one in which the researcher sets up the situation, standardises it and records how participants behave within it. Controlled observations usually take place in a laboratory or a specially arranged room, and the researcher often watches from behind a one-way mirror or through a camera.

The best-known example is Ainsworth’s Strange Situation. In Ainsworth and Bell’s (1970) account, 56 one-year-olds from white, middle-class families were observed in an unfamiliar playroom. Each infant went through the same eight episodes in a fixed order, most lasting three minutes, in which a stranger entered and the mother left and later returned. Observers watched from an adjoining room through a one-way window. Because every baby faced exactly the same sequence, differences in how they reacted could be put down to the babies and their relationships, not to differences in the situation. The procedure is covered in detail in the Mary Ainsworth Strange Situation article.

Strengths of controlled observation:

  • Easy to replicate. A standardised procedure can be repeated by other researchers, in other countries, to check whether the findings hold. The Strange Situation has been used in this way many times.
  • Control of extraneous variables. The setting is the same for everyone, so differences in behaviour are easier to interpret.
  • Easier to record. Cameras and one-way mirrors let several observers watch at once, and recordings can be replayed and coded later.

Weaknesses of controlled observation:

  • Lower ecological validity. People and children may behave differently in an unfamiliar room from the way they behave at home.
  • Participants usually know they are being studied, which can lead to demand characteristics, where people change their behaviour to fit what they think the study is about.
  • The setting may not suit every group. A procedure designed around one culture’s childcare habits may produce different results elsewhere for reasons unrelated to what is being measured.

A controlled observation is not the same as a laboratory experiment. In the Strange Situation nothing is compared between conditions set by the researcher: every infant goes through the same episodes, and the aim is to describe and classify their behaviour. In Bandura’s study, the researcher did create conditions to compare, so the study was an experiment that used a controlled observation as its measure.

Covert vs Overt Observation

The difference between covert and overt observation is whether the people being observed know about it. In a covert observation, participants are unaware that they are being watched or that a study is taking place. In an overt observation, they know they are being observed, and usually have agreed to it.

What Is a Covert Observation?

A covert observation is one where the researcher’s presence or purpose is hidden from the people being studied. The observer might watch from out of sight, use hidden cameras, or join the group while pretending to be an ordinary member.

Rosenhan (1973) is the classic example. Eight sane people gained secret admission to 12 different psychiatric hospitals by complaining that they had been hearing voices. Once admitted, they behaved normally and spent their time writing down observations about the ward, the patients and the staff. They stayed between 7 and 52 days, with an average of 19 days. The pseudopatients were never detected, and the nursing records for three of the pseudopatients treated their writing as part of their illness: “Patient engages in writing behavior” was the daily comment on one of them (Rosenhan, 1973). The patients were more perceptive. During the first three hospitalisations, when accurate counts were kept, 35 of the 118 patients on the admissions ward voiced suspicions that the pseudopatients were not ill.

Notice what made the observation covert. Rosenhan’s pseudopatients did not hide their note-taking for long. The notes were first written secretly, but once it became clear that nobody much cared, they wrote openly in the dayroom (Rosenhan, 1973). What stayed hidden was the purpose: the staff did not know they were being observed for a study.

Strengths of covert observation:

  • No participant reactivity. People who do not know they are being watched have no reason to change their behaviour, so what is recorded is more likely to be genuine.
  • Access to hidden behaviour. Some behaviour, such as how staff treat patients or how a closed group operates, would never be shown to a known researcher.

Weaknesses of covert observation:

  • No informed consent. Participants cannot agree to something they do not know about, and they may be upset to find out later.
  • Privacy. Watching people in settings where they expect privacy is an ethical problem, however valuable the data.
  • Recording is harder. A researcher who must keep a secret cannot always write notes or film, and may have to rely on memory, which can introduce errors.

What Is an Overt Observation?

An overt observation is one in which the people being watched know that they are being observed and, in most cases, have given informed consent. The researcher may sit at the side of a classroom with a clipboard, set up a visible camera, or be introduced to the group as a researcher.

Most observations of young children in nurseries are overt, at least for the adults. Mildred Parten’s (1932) classic study of play was carried out in a nursery school over a whole year, from October 1926 to June 1927, with the observer recording the play of 42 children during their free-play hour. The months from October to January were used partly to develop the technique, and one of the benefits of such a long study is that children become used to the observer’s presence and return to their usual behaviour. The categories that came out of this work are still taught to early years practitioners, and are explained in Mildred Parten’s six stages of play.

Strengths of overt observation:

  • More ethical. Participants can give informed consent, know their right to withdraw, and can be debriefed.
  • Easier to record. The researcher can take notes openly, use checklists and film, which improves accuracy.

Weaknesses of overt observation:

  • Participant reactivity. People who know they are being watched may behave differently, for example by being more helpful, polite or hard-working than usual. This reduces the validity of the findings.
  • Demand characteristics. Participants may try to work out the aim and behave in a way they think the researcher wants, or deliberately the opposite.

Researchers reduce reactivity in overt observations by spending time in the setting before recording begins, so that participants get used to them, and by being as unobtrusive as possible.

What the BPS Code Says About Observing People

The British Psychological Society’s Code of Human Research Ethics allows observation without consent only in limited circumstances. Unless people agree to be observed, observational research is acceptable only in public situations where those observed “would expect to be observed by strangers” (British Psychological Society, 2021). The Code also asks researchers to respect local cultural values and to take care with people who may believe they are unobserved, even in a public place.

Covert data collection falls under the Code’s section on deception. It should happen only where it is essential for the research, where there is no alternative, where the research has strong scientific merit and where there is a plan to manage risk and harm (British Psychological Society, 2021). This is why a covert observation of people in a shopping centre can be acceptable while a covert observation inside a family home is not. For wider ethical issues when research touches sensitive groups, see social sensitivity and research ethics.

Participant vs Non-Participant Observation

The difference between participant and non-participant observation is whether the researcher becomes part of the group being studied. A participant observer joins in with the group’s activities. A non-participant observer stays separate and watches from outside, whether from across the room or from behind a one-way mirror.

What Is Participant Observation?

Participant observation is a technique in which the researcher takes part in the life of the group they are studying, so that they can see its behaviour from the inside. Participant observation can be overt, where the group knows the researcher’s role, or covert, where the researcher pretends to be an ordinary member.

One of the most famous examples in social psychology is When Prophecy Fails (Festinger, Riecken and Schachter, 1956). The researchers joined a small group whose leader, given the pseudonym Marian Keech, predicted that a great flood would destroy much of the world on 21 December 1954 and that believers would be rescued by visitors from space. The researchers posed as believers so they could record what happened when the prediction failed. Instead of abandoning their beliefs, some members became more committed and began trying to convert others. The study became the foundation of cognitive dissonance theory.

It also shows the main problem with participant observation. In a small group, several researchers joining at once become a real part of the group, so their presence and their actions can shape the very events they are recording. Festinger’s study is often criticised on exactly these grounds.

Strengths of participant observation:

  • Insight. Being inside the group lets the researcher understand why people behave as they do, not just what they do.
  • Access. Some groups would only reveal their real behaviour to a member.
  • Rich qualitative data. Participant observation produces detailed descriptions that can reveal things no one thought to measure.

Weaknesses of participant observation:

  • Loss of objectivity. A researcher who spends a long time in a group may come to identify with it and interpret events as members do. This is sometimes called “going native”.
  • The researcher may change the group. Joining in means influencing what happens, which can affect the behaviour being studied.
  • Recording problems. It is hard to take detailed notes while also taking part, so records are often written later from memory.

What Is Non-Participant Observation?

Non-participant observation is a technique in which the researcher remains separate from the group and records behaviour without joining in. Most observations in psychology are non-participant, because they make objective, structured recording much easier.

In Bandura, Ross and Ross (1961), 72 children aged between 37 and 69 months from the Stanford University Nursery School were observed for 20 minutes each. Judges in an adjoining observation room watched through a one-way mirror and rated behaviour in predetermined categories. They did not interact with the children at all, which is what made the observation non-participant.

Strengths of non-participant observation:

  • Objectivity. A researcher who stays outside the group is less likely to be drawn into its point of view.
  • Easier recording. The observer can concentrate fully on watching and coding behaviour, often with a checklist.

Weaknesses of non-participant observation:

  • Less insight. The observer sees behaviour but may misunderstand what it means to the people involved.
  • Reactivity if visible. A stranger sitting at the edge of a room with a clipboard can be very noticeable, unless the observation is also covert.

Structured vs Unstructured Observation

Structured and unstructured observation describe how the behaviour is recorded rather than where or by whom. In a structured observation the researcher decides in advance exactly which behaviours to record and uses a system, such as a tally chart of behavioural categories. In an unstructured observation the researcher writes down everything they see, usually as a continuous narrative.

AQA does not list structured and unstructured among its types of observation, but the distinction sits at the heart of observational design, and it appears in exam questions and other specifications. Bandura’s scoring of set responses in five-second intervals is highly structured. An early years practitioner writing a free description of a child’s morning is using an unstructured observation.

FeatureStructured observationUnstructured observation
What is recordedOnly pre-set behavioural categoriesEverything the observer notices
Type of dataQuantitative: tallies and frequenciesQualitative: detailed descriptions
ReliabilityEasier to check and usually higherHarder to check
RiskImportant behaviour outside the categories is missedObserver bias in what gets written down
Best forTesting a hypothesis about specific behavioursExploring a new area, or small-scale studies

Structured observation produces numbers, which can be summarised and tested statistically. Unstructured observation produces rich description but tends to record whatever stands out most to the observer, which may not be what matters most. The trade-off is the same one explored in quantitative vs qualitative analysis. Many studies combine both: an unstructured pilot phase to discover the behaviours worth recording, then a structured phase to count them.

What Is Observational Design?

Observational design is the set of decisions that turn “watching behaviour” into a study that produces usable, checkable data. The AQA specification names three parts: “behavioural categories; event sampling; time sampling” (AQA, 2015). In practice, designing an observation means answering four questions in order:

  1. What exactly counts as the behaviour? This is the job of behavioural categories.
  2. When will it be recorded? Every time it happens (event sampling) or at set intervals (time sampling).
  3. Who and where? Which participants, chosen how, and in which setting. The choice of participants follows the same rules as any other study, covered in sampling methods in psychology.
  4. How will the records be checked? Usually by inter-observer reliability.

Be careful with the word “sampling” here. Event and time sampling are about when behaviour is recorded. Sampling methods such as random or opportunity sampling are about which people take part. An observation needs both, and exam questions often test whether you can keep them apart.

What Are Behavioural Categories?

Behavioural categories are the specific, observable actions that a researcher records in place of a broad behaviour. They break a general target such as “aggression” or “sociability” into separate behaviours that an observer can see and count, such as “hits with an object” or “kicks”. Creating behavioural categories is a form of operationalisation: defining a variable so precisely that it can be measured.

Bandura, Ross and Ross (1961) give a good example. They did not ask observers to judge whether a child was “aggressive”. They defined separate response categories. Imitative physical aggression included striking the Bobo doll with the mallet, sitting on the doll and punching it in the nose, kicking it and tossing it in the air. Imitative verbal aggression meant the child repeating the model’s phrases, such as “Sock him” or “Pow”. A third category covered repeating the model’s non-aggressive phrases. Because each category named a concrete action, two observers could score the same child and agree.

Parten (1932) did the same for play. Her six categories were unoccupied behaviour, solitary play, onlooker behaviour, parallel play, associative play and cooperative play, each with a written definition and a one-letter code (u, s, o, p, a, c) that the observer wrote on a record form. Solitary play, for instance, was defined as a child playing alone with toys different from those used by children within speaking distance, making no effort to get close to them.

What Makes a Good Behavioural Category?

A good behavioural category is observable, clearly defined, separate from the other categories and, together with them, covers all the behaviour that matters. In more detail:

  • Observable and measurable. “Smiles” can be seen; “feels happy” cannot. Categories should describe actions, not the observer’s interpretation of them.
  • Precisely defined. “Pushes another child with one or both hands” is better than “rough play”, because two observers are less likely to disagree about it.
  • No overlap. Each behaviour should fit only one category. If “pushing” and “physical aggression” are both categories, one push might be recorded twice, or put in different boxes by different observers.
  • Exhaustive. The categories should cover every relevant behaviour, with no gaps, so observers are not left unsure where to record something. An “other” category can catch anything unexpected.
  • Not too many. Observers can only watch and code so much at once. Too many categories reduce accuracy, which is one reason to pilot the checklist first.

The table below shows how a vague target becomes a set of behavioural categories for an imaginary playground study. It is an illustration, not taken from a published study.

Vague targetBehavioural categoriesWhy it works
AggressionHits, pushes, kicks, grabs an object from another child, verbal insultEach is a visible action that can be tallied
FriendlinessSmiles at a peer, offers a toy, invites a peer to join, helps a peerDescribes actions rather than a feeling
Anxiety at separationCries, follows parent to door, searches for parent, refuses to playObservers can agree on what they saw

Event Sampling vs Time Sampling

Event sampling and time sampling are the two ways of deciding when behaviour is recorded in a structured observation. In event sampling the observer records every occurrence of a target behaviour during the observation. In time sampling the observer records what is happening at, or during, fixed time intervals, such as every 30 seconds.

What Is Event Sampling?

Event sampling means the observer decides in advance which behaviours count and then records every instance of them, usually as a tally, for the whole observation period. It suits behaviours that are fairly infrequent and easy to spot when they happen.

Craig and Pepler’s (1998) playground study used episodes of bullying as its unit. Each time an episode happened, it was identified and its details were coded, such as how long it lasted, who was involved and whether an adult or a peer intervened. That is how the researchers could report that episodes lasted 38 seconds on average and that adults stepped in for only 4% of them.

  • Strength: infrequent behaviour is not missed, because every instance is recorded, which suits behaviours that could fall between the intervals of a time sample.
  • Weakness: in a busy, complex situation too much may happen at once for one observer to record everything, which reduces accuracy.
  • Weakness: a simple tally loses context, such as what led up to the behaviour, unless the observer also records details of each event.

What Is Time Sampling?

Time sampling means the observer records behaviour at set intervals, for example noting which behavioural category a child is showing every ten seconds, or observing each child in a group for one minute in turn. It turns a continuous stream of behaviour into a manageable number of records.

Bandura, Ross and Ross (1961) divided each 20-minute session into 5-second intervals with an electric interval timer, which gave 240 response units for every child. Parten (1932) used one-minute samples: the observer watched one child for a minute, recorded the type of play, then moved to the next child on a prearranged list. The order of the list was varied from day to day, and the play hour was split into five-minute sections so that each child was observed equally often at the start, middle and end of the session. That rotation stopped a child’s records being distorted by always being watched at the same time of day. Parten found that twenty one-minute observations per child gave a consistent picture of how that child played: when observations from odd and even days were compared, the correlation was .90.

  • Strength: it reduces the number of observations, so the observer can concentrate and record accurately, and it makes it easy to observe several participants in turn.
  • Strength: it produces data that are easy to compare, such as the proportion of intervals spent in each type of play.
  • Weakness: behaviour that happens between intervals is missed, so the record may not represent the whole session. A brief but important event, such as one push, could go unrecorded.
FeatureEvent samplingTime sampling
What is recordedEvery instance of the target behaviourBehaviour at, or during, set intervals
Best forInfrequent, distinct behavioursFrequent or continuous behaviours, several participants
Main riskToo much to record in a busy settingBehaviour between intervals is missed
ExampleBullying episodes, Craig and Pepler (1998)5-second intervals, Bandura et al. (1961)

Worked Example: Recording a Time Sample

Imagine a student observing one four-year-old during free play, using Parten’s categories and recording every 30 seconds for five minutes. That gives ten records. The student writes the code for whatever the child is doing at each 30-second point:

Interval12345678910
Codeoopppaapsp

The summary is then easy: parallel play in 5 of 10 intervals (50%), onlooker behaviour in 2 (20%), associative play in 2 (20%) and solitary play in 1 (10%). These are illustrative numbers, but they show why time sampling produces quantitative data that can be compared between children or settings. They also show its weakness: if the child briefly joined a cooperative game between interval 6 and interval 7, the record would never show it.

How Do You Check Inter-Observer Reliability?

Inter-observer reliability is the extent to which two or more observers watching the same behaviour record it in the same way. It is checked by having observers record the same behaviour independently, using the same behavioural categories, and then comparing their records statistically. The AQA specification lists inter-observer reliability as one of the two ways of assessing reliability (AQA, 2015).

  1. Agree the behavioural categories and their definitions, and train the observers on them, ideally with practice on video.
  2. Observers watch the same participants at the same time, or the same recording, but record independently without discussing what they see.
  3. Each observer produces their own tally or coding sheet.
  4. Compare the records, usually by correlating the two observers’ totals for each category, or by working out the percentage of records on which they agreed.
  5. If agreement is high, the observation is reliable. If it is low, find the categories causing disagreement, tighten their definitions, retrain and test again.

A correlation of +0.80 or above is the benchmark usually taught at A level for good inter-observer reliability. It is a convention used in teaching materials rather than a figure set in the AQA specification, so the safest exam answer describes it as the usual benchmark.

Worked Example: Correlating Two Observers’ Tallies

Two observers use event sampling to record aggression in the same 20-minute playground session. Their tallies are:

CategoryObserver AObserver B
Hits1211
Pushes78
Kicks33
Verbal insult98
Grabs an object56
Excludes from game21

Because the data are ordinal tallies across categories, a suitable test is Spearman’s rho. Ranking each observer’s totals and applying the formula gives rho = +0.99, well above +0.80, so the two observers’ records agree closely. The steps for ranking and calculating are set out in the guide to Spearman’s rho, which includes a calculator.

Notice what the correlation does and does not show. It shows the two observers ranked the categories the same way. It does not show that they recorded the same individual events, which is why researchers who code each event or interval often use percentage agreement instead.

Percentage Agreement and Cohen’s Kappa

Percentage agreement is the number of records on which the observers agreed divided by the total number of records, multiplied by 100. If two observers code the same 40 time intervals and agree on 36, agreement is 90%. Craig and Pepler (1998) reported exactly this kind of figure, identifying bullying episodes with 90% agreement between raters.

The weakness of percentage agreement is that some agreement happens by chance. If a child spends most of the session in parallel play, two observers who guessed “parallel” every time would agree often without watching closely. Cohen (1960) introduced kappa to correct for this, by comparing the agreement observed with the agreement expected by chance. Kappa runs up to +1, where +1 is perfect agreement. Landis and Koch (1977) described values from 0.61 to 0.80 as “substantial” and above 0.80 as “almost perfect”, though McHugh (2012), writing about health research, argues that any kappa below 0.60 shows inadequate agreement and that little confidence should be placed in results based on it (McHugh, 2012). Kappa goes beyond A level, but knowing that percentage agreement can be inflated by chance is a strong evaluation point.

Inter-Observer Reliability in Classic Studies

Classic observations show how the check works in practice:

  • Bandura, Ross and Ross (1961): half the children were scored independently by a second observer. Because the categories described highly specific, concrete behaviours, agreement was high, with product-moment correlations “in the .90s” (Bandura et al., 1961).
  • Parten (1932): as a check on the objectivity of her categories, three assistants recorded the same child at the same moment as Parten herself, each with a stop-watch and only a printed sheet of definitions to guide them. Parten noted that some disagreements came from unfamiliarity with the codes, which none of them had studied for more than a few minutes, and from the observers watching from slightly different positions.
  • Craig and Pepler (1998): bullying episodes were identified with 90% agreement between raters.

Parten’s account is a useful reminder of why training matters. The disagreements she describes, unfamiliar codes and different viewpoints, are exactly what practice sessions and shared recordings are designed to remove.

What Are the Problems With Observational Research?

The main problems with observational research are observer bias, participant reactivity, the lack of any explanation for the behaviour recorded, and ethical issues around consent and privacy. Most of observational design exists to reduce the first two.

Observer Bias

Observer bias happens when an observer’s expectations, beliefs or knowledge of the hypothesis affect what they notice and how they record it. Hastorf and Cantril (1954) showed how powerful this can be. After a notoriously rough 1951 American football game between Princeton and Dartmouth, students from both universities watched the same film of the game. Princeton students saw the Dartmouth team commit more than twice as many rule infractions as Dartmouth students saw. Both groups watched identical footage; what differed was which team they supported. It is a close relative of confirmation bias, the tendency to notice what fits what we already believe.

Observer bias is reduced by:

  • Clear behavioural categories that leave little room for interpretation.
  • More than one observer, with inter-observer reliability checked.
  • Blind observers who do not know the hypothesis or which condition a participant is in. In Bandura et al. (1961), one or other of the two observers usually did not know which condition a child had been in, although the researchers admit that children who had seen the aggressive model were easy to identify from their distinctive behaviour, which shows the limits of blinding.

Participant Reactivity

Participant reactivity is any change in behaviour caused by people knowing they are being observed. It is the main weakness of overt observation. Covert observation removes it but creates ethical problems, and an observer who spends time in the setting before recording begins allows people to get used to them. Non-participant observers can reduce it further by using one-way mirrors or cameras.

Observation Shows What, Not Why

An observation records what people do but not why they do it, or what they are thinking. A child sitting alone may be content in solitary play or may have been excluded, and the observation alone may not tell the difference. For this reason observations are often combined with interviews or questionnaires. On its own, a naturalistic observation also cannot establish cause and effect, because nothing is manipulated.

Sampling and Generalisation

Many famous observations used small or narrow samples: Bandura’s 72 children all came from one university nursery, Ainsworth and Bell’s 56 infants were from white, middle-class families, and Parten’s 42 children came from a single nursery school. Findings from narrow samples may not generalise to other groups or cultures, a problem discussed in gender and cultural bias in psychology.

Observation in Early Years Settings

Observation is not only a research method. It is also one of the everyday tools of early years practice, and the same techniques appear in nurseries as in psychology studies. Practitioners use narrative observations, which are unstructured records, to describe what a child did and said. They use event samples to track a particular behaviour, such as how often a child joins in with others, and time samples to see how a child spends a session.

The principles are the same. Clear definitions stop two practitioners recording the same behaviour in different ways, and it is worth checking with a colleague whether you would both have written down the same thing. The early years tradition of close observation owes a great deal to pioneers such as Susan Isaacs, whose work centred on careful observation of children’s play.

Practice: Classify the Observation

The quickest way to learn the six types is to classify real scenarios on all three pairs at once. For each scenario below, choose one option from each pair, then check your answer.

Classify the Observation

Eight scenarios. Pick one option from each of the three pairs, then check your answer.

Scenario 1 of 8

Researchers watch children in a school playground from a staffroom window and record every bullying episode. The children are not told.

How to Answer Observation Exam Questions

Exam questions on observational techniques usually ask you to identify the type of observation in a scenario, explain a strength or weakness in context, design an observation, or suggest how to improve its reliability. The skill examiners look for is applying the terms to the study described, not reciting definitions.

Identifying the Type of Observation

Work through the three pairs in turn and point to the evidence in the scenario for each answer:

  1. Setting: did the researcher arrange the situation? If yes, controlled. If behaviour happened where it normally would, naturalistic.
  2. Awareness: did the people being watched know? If not, covert. If they knew, overt.
  3. Role: did the researcher join the group? If yes, participant. If they watched from outside, non-participant.

Then check the recording method. A pre-set checklist means a structured observation; recording every instance of a behaviour means event sampling; recording at intervals means time sampling.

Designing an Observation

A design question, such as "Explain how the psychologist could carry out an observation of aggression in a nursery playground", rewards specific, practical detail. A strong answer covers:

  • Type: for example a naturalistic, non-participant observation, with a reason (behaviour is realistic in the child's usual setting).
  • Behavioural categories: two or three operationalised examples, such as hits, pushes and grabs toys.
  • Sampling: event sampling of each aggressive act, or a time sample such as recording one target child every 30 seconds for ten minutes, with a reason.
  • Reliability: two observers recording independently, then correlating their tallies and looking for +0.80 or above.
  • Ethics: consent from parents and the setting, and how children's identities will be protected.

Evaluating in Context

Evaluation points gain marks when they are tied to the scenario. "Covert observations are unethical" is generic. "Because the shoppers did not know they were being watched, they could not give informed consent, but the observation took place in a public space where people would expect to be seen by strangers, so it may still be acceptable under the BPS Code" is applied, balanced and specific.

Conclusion

Observational techniques let psychologists study what people do rather than what they say they do. Every observation combines one option from each of three pairs: naturalistic or controlled, covert or overt, and participant or non-participant. Each choice trades something: realism against control, honest behaviour against ethics, and insight against objectivity. Good observational design reduces those costs by turning behaviour into clear behavioural categories, recording it with event or time sampling, and checking inter-observer reliability. Classic studies from Parten's nursery to Bandura's Bobo doll show how carefully this has to be done, and Rosenhan's hospitals show how much can be learned when it is.

Frequently Asked Questions

What are the 6 types of observation in psychology?

The six types are naturalistic, controlled, covert, overt, participant and non-participant observation. They form three pairs, and a single study is always one of each pair. For example, watching children in their own nursery from the side of the room, with parents' consent, is naturalistic, overt and non-participant.

What is an example of a naturalistic observation?

A well-known example is Craig and Pepler's study of bullying, in which researchers observed children during ordinary school playtimes. They found bullying happened about once every seven minutes, but adults intervened in only 4% of episodes. Other examples include recording how shoppers queue or how toddlers share toys at nursery.

What is the difference between covert and overt observation?

In covert observation the people being watched do not know they are being studied; in overt observation they do. Covert observation gives more natural behaviour but raises ethical problems because there is no informed consent. Overt observation is more ethical but people may change their behaviour because they know they are being watched.

What are the advantages and disadvantages of observation in psychology?

The main advantage is that observation records what people actually do rather than what they report, often in realistic settings. The main disadvantages are observer bias, people behaving differently when watched, and the fact that observation shows what happened but not why. Naturalistic observations also cannot show cause and effect.

What is the difference between event sampling and time sampling?

Event sampling records every occurrence of a chosen behaviour during the observation. Time sampling records what is happening at fixed intervals, such as every 30 seconds. Event sampling suits rare, distinct behaviours; time sampling suits frequent behaviours or watching several people in turn, but can miss things that happen between intervals.

What are behavioural categories in psychology?

Behavioural categories are the specific, observable actions a researcher records instead of a broad behaviour. "Aggression" might be broken down into hitting, kicking and pushing. They should be clearly defined, should not overlap, and should cover all the behaviour of interest, so different observers can record the same thing.

What is inter-observer reliability?

Inter-observer reliability is the extent to which different observers agree when recording the same behaviour. They observe independently using the same categories, then their records are compared, usually with a correlation. A correlation of +0.80 or more is the benchmark usually taught at A level.

What is a structured observation?

A structured observation uses a pre-planned system, such as a checklist of behavioural categories and a sampling method, to record only the behaviours the researcher has decided on in advance. It produces numerical data that are easy to compare. An unstructured observation, by contrast, records everything the observer sees as a written description.

What is time sampling in early childhood education?

In early childhood education, time sampling means a practitioner records what a child is doing at regular intervals, for example every five minutes across a morning. It builds a picture of how the child spends their time, such as where they play and who with, without the practitioner having to watch continuously.

References

  • Ainsworth, M. D. S., & Bell, S. M. (1970). Attachment, exploration, and separation: Illustrated by the behavior of one-year-olds in a strange situation. Child Development, 41(1), 49–67.
  • AQA. (2015). AS and A-level Psychology specification (7181, 7182). AQA.
  • Bandura, A., Ross, D., & Ross, S. A. (1961). Transmission of aggression through imitation of aggressive models. Journal of Abnormal and Social Psychology, 63(3), 575–582.
  • British Psychological Society. (2021). BPS code of human research ethics. British Psychological Society.
  • Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46.
  • Craig, W. M., & Pepler, D. J. (1998). Observations of bullying and victimization in the school yard. Canadian Journal of School Psychology, 13(2), 41–59.
  • Craig, W. M., Pepler, D., & Atlas, R. (2000). Observations of bullying in the playground and in the classroom. School Psychology International, 21(1), 22–36.
  • Festinger, L., Riecken, H. W., & Schachter, S. (1956). When prophecy fails. University of Minnesota Press.
  • Hastorf, A. H., & Cantril, H. (1954). They saw a game: A case study. Journal of Abnormal and Social Psychology, 49(1), 129–134.
  • Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174.
  • McHugh, M. L. (2012). Interrater reliability: The kappa statistic. Biochemia Medica, 22(3), 276–282.
  • Parten, M. B. (1932). Social participation among pre-school children. Journal of Abnormal and Social Psychology, 27(3), 243–269.
  • Rosenhan, D. L. (1973). On being sane in insane places. Science, 179(4070), 250–258.

Further Reading and Research

Recommended Articles

Suggested Books

  • Coolican, H. (2019). Research Methods and Statistics in Psychology (7th ed.). Routledge.
    • The standard research methods textbook for A level and undergraduate psychology, with a full chapter on observational methods, coding and inter-observer reliability.
  • Bakeman, R., & Quera, V. (2011). Sequential Analysis and Observational Methods for the Behavioral Sciences. Cambridge University Press.
    • An advanced guide to designing coding schemes, recording behaviour in time and measuring observer agreement, for readers going beyond A level.
  • Festinger, L., Riecken, H. W., & Schachter, S. (1956). When Prophecy Fails. University of Minnesota Press.
    • The original account of a covert participant observation, readable as a narrative and a vivid case for discussing its ethics.

Recommended Websites

  • AQA A-level Psychology (7182)
    • The official specification, past papers and mark schemes, including the research methods content on observational techniques and design.
  • Classics in the History of Psychology
    • Free full texts of classic papers, including Bandura, Ross and Ross (1961), where you can read exactly how the observation was designed.
  • British Psychological Society
    • The Code of Human Research Ethics, which sets out when observing people without their consent is acceptable.

Kathy Brodie

Kathy Brodie is an Early Years Professional, Trainer and Author of multiple books on Early Years Education and Child Development. She is the founder of Early Years TV and the Early Years Summit.

Kathy's Author Profile
Kathy Brodie

To cite this article please use:

Early Years TV Observational Techniques in Psychology: Types and Design. Available at: https://www.earlyyears.tv/observational-techniques-psychology-types-design/ (Accessed: 2 October 2026).