Reliability in Psychology: Types, Assessment and Improvement

When two psychiatrists independently assessed the same patients in the DSM-5 field trials, their agreement on major depressive disorder reached a kappa of just 0.28, a level of reliability the researchers themselves rated questionable.
Key Takeaways
- What is reliability? Reliability is consistency: a reliable measure gives the same result when it is repeated, or when a different researcher uses it.
- How is it assessed? AQA names two ways, test-retest and inter-observer. Both compare two sets of results, and a correlation of +0.80 or above is taken to show reliability.
- How is it improved? By tightening the procedure: standardised instructions, operationalised behavioural categories, trained observers, clearer questions and a pilot study.
Reliability is one of the few ideas that runs through every part of the A level psychology course. It appears in research methods, where you have to say how it is assessed and improved, and again in Paper 3, where the reliability of diagnosing schizophrenia or depression is a standard evaluation point. The AQA specification requires “reliability across all methods of investigation”, which means you need to be able to apply it to experiments, observations, questionnaires, interviews, content analysis and case studies alike (AQA, 2021).
The good news is that reliability comes down to one simple question: if we did this again, would we get the same answer? Everything else, from test-retest checks to Cohen’s kappa, is a way of putting a number on that answer. Most of those numbers are correlations, so if you are comfortable with a correlation coefficient you already have most of the maths you need.
This guide explains what reliability means in psychology, the difference between internal and external reliability, how each type is assessed and improved, how reliability differs from validity, and why it matters for real decisions such as psychiatric diagnosis. It uses real data throughout, including the inter-observer figures from Bandura’s Bobo doll study and a genuine AQA exam question on improving the reliability of a content analysis. There is also a calculator for checking reliability from your own data and a quiz to test whether you can tell the types apart.
What Is Reliability in Psychology?
Reliability in psychology is the consistency of a measure or a study. A reliable measure produces the same results each time it is used under the same conditions, and a reliable study produces the same findings when it is repeated. If a personality test puts you in one category today and a different one next week, even though nothing about you has changed, the test is unreliable.
A useful everyday comparison is a set of bathroom scales. If you step on three times in a row and get 60 kg, 64 kg and 57 kg, the scales are unreliable, because your weight has not changed in thirty seconds. If you get 60 kg every time, the scales are reliable. Notice that reliable does not mean correct: scales that read 60 kg every time when you actually weigh 63 kg are perfectly reliable and consistently wrong. That second problem is validity, which is covered later in this guide.
Reliability Is About Measurement Error
Reliability matters because every psychological measurement contains some error. A score on a memory test reflects the person’s real memory ability, plus a little noise from tiredness, distraction, a lucky guess or an ambiguous question. The more of the score that comes from real differences between people, and the less from noise, the more reliable the measure is.
That noise is not harmless. Hedge, Powell and Sumner (2018) showed that when a measure has low reliability, correlations between it and anything else shrink and become hard to replicate, which can lead researchers to wrong conclusions about how psychological abilities relate to each other or to the brain. An unreliable measure does not just add a bit of fuzz; it can hide a relationship that is really there.
Reliability of Measures and Reliability of Studies
Reliability applies at two levels, and exam answers are stronger when you say which one you mean. The first is the reliability of a measure: a questionnaire, a test, a set of behavioural categories or a diagnostic interview. The second is the reliability of a study: whether the whole procedure, repeated with new participants, would produce the same findings. The second depends heavily on the first, but also on standardisation, the controlled, identical procedure every participant goes through, which is why tightly controlled experiments are usually described as more reliable than unstructured methods.

Types of Reliability: Internal and External
The two types of reliability are internal reliability and external reliability. Internal reliability is consistency within a measure: whether all the parts of a test or questionnaire measure the same thing. External reliability is consistency of a measure over time or between people: whether it gives the same result on a different occasion, or when used by a different researcher.
| Type | Question it answers | How it is assessed |
|---|---|---|
| Internal reliability | Do all the items in this measure agree with each other? | Split-half method; Cronbach’s alpha; item-total correlations |
| External reliability over time | Does the measure give the same result on a second occasion? | Test-retest method |
| External reliability between people | Do two researchers using the measure get the same result? | Inter-observer (inter-rater) reliability |
Internal Reliability
Internal reliability is the extent to which a measure is consistent within itself. Take a 20-item questionnaire on self-esteem. If it is internally reliable, someone with high self-esteem should score high on most of the items, and someone with low self-esteem should score low on most of them. If one item behaves differently, for example an item about sleep that has crept into a self-esteem scale, it lowers internal reliability because it is measuring something else.
Internal reliability only makes sense for measures with several parts that are supposed to measure the same thing, so it applies mainly to questionnaires, psychometric tests and rating scales. It is assessed by the split-half method or by Cronbach’s alpha, both explained below.
External Reliability
External reliability is the extent to which a measure gives consistent results from one use to the next. It comes in two forms: consistency over time, assessed by the test-retest method, and consistency between researchers, assessed by inter-observer reliability. External reliability applies to almost every method in psychology, which is why these two checks are the ones exam boards focus on.
What Each Exam Board Requires
The three main A level boards cover reliability in slightly different ways, so check which list applies to you. AQA names two ways of assessing reliability, test-retest and inter-observer, plus improving reliability (AQA specification, section 4.2.3). OCR lists five types by name. Edexcel lists reliability alongside objectivity and validity, and applies it specifically to the diagnosis of mental disorders.
| Board | What the specification says |
|---|---|
| AQA (7182) | Reliability across all methods of investigation; ways of assessing reliability: test-retest and inter-observer; improving reliability. Also reliability in the diagnosis of schizophrenia. |
| OCR (H567) | Reliability: internal, external, inter-rater, test-retest and split-half. |
| Edexcel (9PS0) | Objectivity, reliability and validity; reliability and validity of diagnoses using DSM and ICD. |
Sources: AQA (2021), OCR (2026) and Pearson Edexcel (2026). Whichever board you sit, learning all five types costs little extra and lets you evaluate any study you meet.
Test-Retest Reliability
Test-retest reliability is the consistency of a measure over time. It is assessed by giving the same test or questionnaire to the same people on two separate occasions, then correlating the two sets of scores. If people who scored high the first time also score high the second time, and low scorers stay low, the measure has good test-retest reliability.
How to Carry Out a Test-Retest Check
- Give the measure to a group of participants under standardised conditions.
- Wait long enough that they cannot simply remember their answers, but not so long that the thing being measured has genuinely changed.
- Give exactly the same measure to the same participants under the same conditions.
- Correlate each person’s first score with their second score.
- A correlation of +0.80 or above indicates good test-retest reliability.
Choosing the Gap Between Tests
The time between the two tests is the hardest decision in a test-retest check, because both short and long gaps cause problems. Too short, and participants remember their answers or benefit from practice, which makes the measure look more reliable than it is. Too long, and real change creeps in, which makes a good measure look unreliable.
Researchers usually pick a gap that suits what they are measuring. Hedge et al. (2018), testing seven well-known attention and inhibition tasks, used two sessions three weeks apart at the same time of day, and kept the task order identical in both sessions so that order effects could not masquerade as unreliability. For something as stable as adult personality, much longer gaps are possible. Roberts and DelVecchio (2000) combined 3,217 test-retest correlations from 152 longitudinal studies and found that personality consistency rose from .31 in childhood to .54 in the college years, .64 at age 30 and about .74 between ages 50 and 70, with the time interval held constant at 6.7 years. They also found that the longer the interval, the lower the consistency, which is exactly what you would expect when a gap allows genuine change.
That finding contains a useful exam point: a low test-retest correlation does not always mean a bad measure. It may mean the thing being measured has changed, which is why test-retest is best suited to characteristics that are meant to be stable, such as personality traits or intelligence, and poorly suited to states that change from hour to hour, such as mood.
Test-Retest Reliability in Personality Testing
Personality tests are where most people meet test-retest reliability in everyday life, usually without realising it. If an online test gives you a different four-letter type every time you take it, it is failing a test-retest check. We look at this problem in detail in our guide to why MBTI results keep changing, which is a good example of how a measure can be popular and still inconsistent.
Strengths and Limitations of Test-Retest
| Strengths | Limitations |
|---|---|
| Simple to carry out and easy to interpret | Practice and memory effects can inflate the correlation |
| Works for almost any questionnaire, test or scale | Genuine change between tests can deflate it |
| Tests the measure in the way it will actually be used | Participants may drop out before the retest, and those who return may differ from those who do not |
Inter-Observer Reliability
Inter-observer reliability is the extent to which two or more observers agree when they record the same behaviour. It is assessed by having observers watch the same events independently, using the same behavioural categories, then comparing their records. When the records come from ratings or interviews rather than observation, the same idea is called inter-rater reliability, and in content analysis it is sometimes called inter-coder reliability.
Inter-observer reliability matters because observation depends on human judgement. One observer might count a push as aggression while another counts it as play. Unless the categories are defined precisely, the data reflect the observer as much as the participant, and a different observer repeating the study would get different results.
How to Assess Inter-Observer Reliability
- Agree operationalised behavioural categories, each defined by observable actions, for example “hits the doll with the mallet” rather than “is aggressive”.
- Train the observers together and run a short pilot on practice footage.
- Have at least two observers watch the same participants at the same time, or the same recording, without conferring.
- Each observer tallies behaviour in the same way, often using time sampling or event sampling.
- Correlate the observers’ tallies. A correlation of +0.80 or above shows good inter-observer reliability.
A Real Example: Bandura’s Bobo Doll Study
Bandura, Ross and Ross (1961) is one of the clearest published examples of inter-observer reliability in action, and it is worth knowing in detail. Before the experiment, the children were rated for how aggressive they normally were, so that they could be matched across conditions. Fifty-one children were rated independently by two judges, and the reliability of the composite aggression score, calculated with a Pearson correlation, was .89.
In the main observation, each child’s 20-minute session was divided into 5-second intervals, giving 240 response units per child, a form of time sampling. Half the children were scored independently by a second observer. The paper reports that the behaviours were “highly specific concrete classes of behavior” and that the interscorer reliabilities were in the .90s (Bandura et al., 1961). Both figures clear the +0.80 benchmark comfortably, and the reason is instructive: the categories were things like striking the doll with a mallet or sitting on it and punching its nose, which leave little room for disagreement. You can read more about the study in our guide to Bandura’s social learning theory and the Bobo doll.
The paper also shows a limitation that examiners like. One of the two observers usually did not know which condition a child had been in, but Bandura et al. (1961) admit that children who had seen the aggressive model could be “readily identified” from their distinctive behaviour. High agreement between observers does not rule out both of them being influenced by the same expectations.
Worked Example: Checking Two Observers’ Tallies
Here is a worked example of the kind you may be asked to interpret. Two observers watched eight children in a nursery playground and each tallied acts of aggression using the same behavioural categories.
| Child | A | B | C | D | E | F | G | H |
|---|---|---|---|---|---|---|---|---|
| Observer 1 | 12 | 7 | 15 | 3 | 9 | 11 | 5 | 14 |
| Observer 2 | 11 | 8 | 14 | 4 | 9 | 13 | 5 | 12 |
The two observers rarely give exactly the same number, but that is not what reliability requires. What matters is whether they rank the children in the same order. Both means are 9.5. Working through Pearson’s formula, the sum of the cross-products is 105, and the sums of squared deviations are 128 for Observer 1 and 94 for Observer 2, so r = 105 divided by the square root of (128 × 94), which is r = +0.96. That is well above +0.80, so the observers have high inter-observer reliability. Because tallies are often treated as ordinal data, you could also use Spearman’s rho, which gives +0.93 here and leads to the same conclusion.
How Is Reliability Measured? Correlation, Agreement and Kappa
Reliability is measured by comparing two sets of results and expressing how closely they agree as a single number. For scores and tallies that number is usually a correlation coefficient; for categories it is percentage agreement or Cohen’s kappa; and for the internal consistency of a questionnaire it is the split-half correlation or Cronbach’s alpha.
The +0.80 Rule
At A level, a correlation of +0.80 or above is the accepted sign of good reliability. You will see it in AQA’s own mark schemes: the June 2023 Paper 1 mark scheme, in a question on improving the reliability of a content analysis, credits the point that the two researchers’ results are compared and that “+0.8 or above would indicate reliability” (AQA, 2023). Quote the figure in any answer about assessing reliability; it turns a vague “compare the results” into a precise, creditworthy test.
Two details are worth adding for top marks. The correlation must be positive, because a strong negative correlation between two observers would mean they systematically disagree. And a correlation only shows that people are ranked in the same order: if one observer consistently counts two more acts of aggression than the other, the correlation can still be high. Researchers who need exact agreement use the intraclass correlation (ICC) instead, for which Koo and Li (2016) suggest values below 0.5 indicate poor reliability, 0.5 to 0.75 moderate, 0.75 to 0.9 good and above 0.9 excellent.
Percentage Agreement
Percentage agreement is the simplest measure of inter-rater reliability for categories: the number of times two coders agree, divided by the number of decisions, multiplied by 100. McHugh (2012) notes that many texts recommend 80% agreement as the minimum acceptable level. Its weakness is that it ignores luck. If a behaviour happens in 90% of intervals, two observers who both simply write “yes” every time will agree 90% of the time without watching at all.
Cohen’s Kappa
Cohen’s kappa is a measure of agreement between two raters that corrects for the agreement expected by chance. Cohen (1960) designed it for exactly the problem above. Kappa is calculated as the observed agreement minus chance agreement, divided by one minus chance agreement. A kappa of 1 means perfect agreement, 0 means no better than chance, and negative values mean worse than chance.
Here is a worked example. Two observers watched a child for 20 intervals and coded each one as on-task (T) or off-task (O). They agreed on 18 of the 20 intervals, which is 90% agreement and looks excellent. But each observer coded 16 intervals as on-task and 4 as off-task, so the agreement you would expect by chance alone is (0.8 × 0.8) + (0.2 × 0.2) = 0.68. Kappa is therefore (0.90 − 0.68) ÷ (1 − 0.68) = 0.69. The agreement is real, but it is noticeably less impressive once luck is removed, because most of the intervals were the easy, common category.
Landis and Koch (1977) proposed the benchmarks most often used to describe kappa values, and they are widely quoted in research, including by the US Agency for Healthcare Research and Quality.
| Kappa | Strength of agreement |
|---|---|
| Below 0 | Poor (less than chance) |
| 0.00-0.20 | Slight |
| 0.21-0.40 | Fair |
| 0.41-0.60 | Moderate |
| 0.61-0.80 | Substantial |
| 0.81-1.00 | Almost perfect |
These labels are generous. McHugh (2012), writing for health researchers, argues that any kappa below 0.60 indicates inadequate agreement, which is a stricter standard than the “moderate” label Landis and Koch give values between 0.41 and 0.60. On either scheme, the 0.69 in the worked example is acceptable but not outstanding.
The Split-Half Method
The split-half method assesses internal reliability by dividing a test into two halves and correlating people’s scores on each half. The usual split is odd-numbered items against even-numbered items, which avoids comparing the easy start of a test with the harder end, or the fresh start with the tired finish. If the two halves correlate strongly, the items are measuring the same thing.
There is one technical wrinkle. Each half is only half as long as the real test, and shorter tests are less reliable, so the raw split-half correlation underestimates the reliability of the full test. The Spearman-Brown formula, published independently by Spearman (1910) and Brown (1910) in the same issue of the British Journal of Psychology, corrects this: full-test reliability = 2r ÷ (1 + r). A split-half correlation of 0.70 therefore corresponds to a full-test reliability of 0.82.
Cronbach’s Alpha
Cronbach’s alpha is the most widely used measure of internal reliability. Cronbach (1951) showed that alpha is, in effect, the average of all the possible split-half coefficients, so it does not depend on which way the test happens to be split. Alpha runs from 0 to 1, with higher values meaning the items hang together more closely.
Tavakol and Dennick (2011) note that reported acceptable values of alpha range from 0.70 to 0.95, and warn that a very high alpha, above 0.90, may mean some items are redundant, asking the same question in a different form. They also point out that alpha rises with test length, so a long questionnaire can reach a high alpha even if it is not measuring a single, unified idea. A high alpha is evidence of consistency, not proof that a scale measures one thing.
Reliability Calculator: Check Your Own Data
This calculator checks test-retest or inter-observer reliability from two sets of data. Choose scores mode for tallies or test scores, which gives Pearson’s r and Spearman’s rho against the +0.80 benchmark, or categories mode for coded data, which gives percentage agreement and Cohen’s kappa. The worked example button loads the two examples above so you can check the working.
Reliability Calculator
Check test-retest or inter-observer reliability. Enter the two sets in the same order, separated by commas, spaces or new lines.
| Statistic | Used for | Usual benchmark |
|---|---|---|
| Correlation (r or rho) | Test-retest and inter-observer scores | +0.80 or above (AQA mark scheme) |
| Percentage agreement | Two coders using categories | 80% or above (McHugh, 2012) |
| Cohen’s kappa | Two coders, corrected for chance | 0.61-0.80 substantial, 0.81-1.00 almost perfect (Landis and Koch, 1977) |
| Intraclass correlation (ICC) | Test-retest and rater agreement in research | 0.75-0.90 good, above 0.90 excellent (Koo and Li, 2016) |
| Cronbach’s alpha | Internal reliability of a questionnaire | 0.70 to 0.95 (Tavakol and Dennick, 2011) |
How to Improve Reliability in Psychology
Reliability is improved by reducing the room for things to vary that should not vary: the procedure, the instructions, the wording of questions and the judgements of the researchers. The specific fix depends on the research method, and exam questions almost always name a method, so tailor your answer to it.
| Method | Main threat to reliability | How to improve it |
|---|---|---|
| Experiments | Procedure varies between participants | Standardised procedures and instructions; control of extraneous variables |
| Observations | Observers interpret behaviour differently | Operationalised behavioural categories; observer training; pilot study; inter-observer check |
| Questionnaires | Ambiguous or leading questions | Clear, closed questions; remove or rewrite items that do not fit; test-retest check |
| Interviews | Different interviewers ask questions differently | Structured interviews; same trained interviewer; recordings for double-coding |
| Content analysis | Coders apply categories differently | Agreed, operationalised coding categories; coder training; trial coding; inter-rater check |
| Case studies | Unique, unrepeatable data and one researcher's interpretation | Several sources of data; more than one researcher interpreting the data |
Improving Reliability in Experiments
Experiments are improved by standardisation: every participant receives the same instructions, the same materials and the same procedure in the same conditions. Loftus and Palmer (1974) is a good illustration. Participants watched the same film clips and answered the same questionnaire, with only the verb in the critical question changing, so any difference in speed estimates can be put down to that verb rather than to how the study was run. Standardisation is also what makes it possible for someone else to repeat the study. The choice of experimental design matters too: in a repeated measures design, counterbalancing stops order effects adding inconsistent noise to one condition.
Improving Reliability in Observations
Observations are improved mainly through better behavioural categories. Each category should be operationalised, meaning defined in terms of specific, observable actions, and the categories should not overlap, so that every act fits in only one place. Observers should then be trained together and practise on recordings until their agreement reaches +0.80, before the real data are collected. Structured procedures such as Ainsworth's Strange Situation, which runs through a fixed sequence of episodes, show how standardising the situation itself also helps observers agree.
Improving Reliability in Questionnaires and Interviews
Questionnaires are improved by rewriting or removing items. Ambiguous questions are answered differently on different days, so replacing open questions with closed ones, or vague wording with specific wording, raises test-retest reliability. Items that do not correlate with the rest of the scale can be removed to raise internal reliability. Interviews become more reliable when they are structured, with the same questions asked in the same order, and when the same interviewer, or interviewers trained to the same script, carry them out.
Improving Reliability in Content Analysis and Case Studies
Content analysis faces the same problem as observation, with coders in place of observers, so the same solutions apply: agree clearly defined coding categories, train the coders, trial the categories on a small sample of material and check inter-rater reliability. Case studies are the hardest method to make reliable, because the case is usually unique and cannot be repeated. The best that can be done is to use several sources of evidence and have more than one researcher interpret the data, so that the conclusions do not rest on one person's reading.
Reliability vs Validity: What Is the Difference?
Reliability is about consistency and validity is about accuracy. A reliable measure gives the same result every time; a valid measure measures what it claims to measure. The two are linked in one direction only: a measure cannot be valid if it is unreliable, but it can be reliable without being valid.
| Reliability | Validity | |
|---|---|---|
| Core question | Is it consistent? | Is it accurate? |
| Everyday example | A clock that is always exactly ten minutes fast | A clock that shows the right time |
| Assessed by | Test-retest, inter-observer, split-half, Cronbach's alpha | Face, concurrent, ecological and temporal validity checks |
| Can exist without the other? | Yes: consistently wrong is still reliable | No: an inconsistent measure cannot be accurate |
The clock example is the easiest way to remember it. A clock that is always ten minutes fast is perfectly reliable, because it is wrong by the same amount every time, but it is not valid. A clock that runs randomly fast and slow is neither. Only a clock that is consistently right is both.
The Reliability Paradox
There is a twist that most revision guides leave out. Reliability in the sense of "the effect always turns up" is not the same as reliability in the sense of "the measure ranks individuals consistently". Hedge et al. (2018) tested seven classic tasks, including the Stroop and Eriksen flanker tasks, whose effects are some of the most replicable in psychology. Their test-retest reliabilities ranged from 0 to .82 and were surprisingly low for most tasks. The reason was not sloppy measurement but low variation between people: the tasks work so reliably precisely because almost everyone shows the effect to a similar degree, which leaves little difference between individuals for a test-retest correlation to detect. A task can be excellent for an experiment and poor for measuring individual differences, which is a strong evaluation point if you are ever asked about the reliability of a cognitive measure.
Reliability in Diagnosis: The DSM-5 Field Trials
Reliability of diagnosis is the extent to which different clinicians give the same diagnosis to the same patient. It is a required topic for AQA's schizophrenia option and Edexcel's clinical psychology unit, and the best modern evidence comes from the field trials run before the publication of DSM-5, the American Psychiatric Association's diagnostic manual.
In those trials, patients at eleven academic centres in the United States and Canada were randomly assigned to two clinicians, who each interviewed them on separate occasions without knowing any previous diagnosis. Agreement was measured with an intraclass kappa (Regier et al., 2013). The researchers classed kappa values of 0.60-0.79 as very good, 0.40-0.59 as good, 0.20-0.39 as questionable and below 0.20 as unacceptable.
| DSM-5 diagnosis | Kappa | Rating in the trial |
|---|---|---|
| Major neurocognitive disorder | 0.78 | Very good |
| Posttraumatic stress disorder | 0.67 | Very good |
| Bipolar I disorder | 0.56 | Good |
| Borderline personality disorder | 0.54 | Good |
| Schizophrenia | 0.46 | Good |
| Major depressive disorder | 0.28 | Questionable |
| Generalised anxiety disorder | 0.20 | Questionable |
| Mixed anxiety-depressive disorder | −0.004 | Unacceptable |
Source: Regier et al. (2013), pooled or single-site intraclass kappa values. Two points stand out. First, the ratings used in the trial are more lenient than the Landis and Koch benchmarks: schizophrenia's 0.46 is "good" here but only "moderate" on the Landis and Koch scale. Second, some of the most commonly diagnosed conditions were among the least reliable. Two clinicians assessing the same patient often disagreed about whether they had major depression, and the proposed mixed anxiety-depressive disorder produced agreement no better than chance.
For an essay on schizophrenia, the figure of 0.46 is a precise, sourced way to make the point that diagnosis is only moderately reliable, and it links neatly to the wider problem of symptom overlap. It also shows why reliability comes before validity in diagnosis: if two psychiatrists cannot agree whether a patient has a disorder, the diagnosis cannot be accurately identifying a real underlying condition. The same logic applies to the older approaches covered in our guide to defining abnormality, where judgements about what counts as abnormal can vary between clinicians.
Reliability, Replication and the Replication Crisis
Replication is the reliability of a study as a whole: whether repeating the study with new participants produces the same findings. It is closely related to external reliability, and it is also one of the features of science on the AQA specification, where it appears as replicability. A finding that cannot be replicated may have been produced by chance, by an unusual sample or by a procedure that was never properly standardised.
The largest test of replication in psychology found it wanting. The Open Science Collaboration (2015) repeated 100 experimental and correlational studies published in three psychology journals. Ninety-seven per cent of the original studies had statistically significant results, but only 36% of the replications did, and the replication effects were on average half the size of the originals. Replication success was better predicted by the strength of the original evidence than by who carried out the replication. That finding is now a standard reference point for any discussion of statistical significance and the risk of false positives.
The lesson for reliability is that a single study, however well designed, is weak evidence on its own. Standardised procedures and reliable measures make replication possible; only repeated replication shows that a finding is dependable.
Quiz: Spot the Reliability Check
The quickest way to lose marks on reliability is to mix up the types, or to describe a validity check as a reliability check. Work through these eight scenarios and pick the type of reliability each one tests.
Spot the Reliability Check
Eight research scenarios. Decide which kind of reliability each one checks, or whether it is really about validity.
Scenario 1 of 8
A researcher gives 40 students an anxiety questionnaire in September and the same questionnaire three weeks later, then correlates the two sets of scores.
How to Answer Reliability Exam Questions
Reliability questions reward precision: name the specific check, say what is compared, and give the benchmark. A real example shows how. The June 2023 AQA Paper 1 asked students to "Explain how the reliability of this content analysis could be improved" for 4 marks, about two researchers analysing diary extracts for signs of depression (AQA, 2023).
The mark scheme credits points such as these: the researchers should agree and operationalise their coding categories, making them mutually exclusive; they should be trained in using those categories; they should trial the system on a small number of diary extracts; and they should check inter-rater reliability by comparing their results, with +0.8 or above indicating reliability. The mark scheme also notes that no credit is given for simply adding another researcher. A top-band answer strings these together into a clear sequence applied to the diary extracts, rather than listing general points about reliability.
Common Mistakes to Avoid
- Confusing reliability with validity. "The study is unreliable because it was done in a lab" is a validity point. Reliability is only about consistency.
- Saying "repeat the study" and stopping. Say what is repeated, what is compared, how it is compared (a correlation) and what counts as reliable (+0.80).
- Forgetting the second observer is independent. Observers who confer while recording will agree, but that agreement proves nothing.
- Using the everyday meaning. In eyewitness testimony, "unreliable" often means inaccurate. In research methods, reliability means consistency, so be precise about which you mean.
- Giving a generic answer. If the question describes a questionnaire, talk about questions; if it describes an observation, talk about behavioural categories.
Conclusion
Reliability is consistency. Internal reliability asks whether the parts of a measure agree, and is checked with the split-half method or Cronbach's alpha. External reliability asks whether a measure gives the same results over time, checked by test-retest, and between researchers, checked by inter-observer reliability. For A level, the benchmark is a correlation of +0.80 or above, and the ways to improve reliability are standardised procedures, operationalised categories, trained researchers, clearer questions and pilot studies.
Reliability is necessary for validity but never enough on its own, and it has real consequences outside the classroom, from whether two psychiatrists give a patient the same diagnosis to whether a published finding can be trusted. Learn the checks, learn the benchmark, and always tailor your answer to the method in front of you.
Frequently Asked Questions
What is reliability in psychology?
Reliability in psychology means consistency. A reliable test, questionnaire or observation gives the same results when it is used again in the same conditions or by a different researcher. A study is reliable if repeating it produces the same findings. Reliability says nothing about accuracy, which is a separate question called validity.
What is the difference between reliability and validity in psychology?
Reliability is whether a measure is consistent, and validity is whether it measures what it claims to measure. A set of scales that always reads 3 kg too heavy is reliable but not valid. Validity depends on reliability, because a measure that changes randomly cannot be accurate, but reliability does not guarantee validity.
What is test-retest reliability?
Test-retest reliability is the stability of a measure over time. The same people complete the same test on two occasions, and their two sets of scores are correlated. A strong positive correlation, conventionally +0.80 or higher, shows the test gives consistent results. The gap between sessions must be long enough to prevent remembered answers but short enough that nothing real has changed.
What is inter-rater reliability in psychology?
Inter-rater reliability is the level of agreement between two or more people who independently rate, score or code the same material. In observational research it is called inter-observer reliability. It is judged by correlating the raters' scores, or for categories by percentage agreement or Cohen's kappa. High inter-rater reliability means the result depends on the participant, not on who did the rating.
What is the difference between internal and external reliability?
Internal reliability concerns consistency inside a single measure: do all its items point the same way? External reliability concerns consistency across uses: does the measure give matching results on another day or with another researcher? Split-half and Cronbach's alpha test internal reliability; test-retest and inter-observer checks test external reliability.
How can you improve the reliability of a questionnaire in psychology?
- Pilot the questionnaire and look for questions people misread or answer inconsistently.
- Rewrite ambiguous items and use closed questions or fixed rating scales where possible.
- Drop items that correlate poorly with the total score.
- Give the same standardised instructions to everyone.
- Run a test-retest check and aim for +0.80 or above.
What is split-half reliability?
Split-half reliability is a way of checking internal consistency by dividing one test into two halves, usually odd against even items, and correlating each person's score on the two halves. Because each half is shorter than the full test, the result is adjusted upwards with the Spearman-Brown formula to estimate the reliability of the whole test.
How do you interpret Cronbach's alpha?
Cronbach's alpha is read on a scale from 0 to 1, where higher values mean the items of a scale are more consistent with one another. Values from about 0.70 up to 0.95 are usually treated as acceptable. Values above 0.90 can signal redundant items, and because alpha grows with the number of items, a long scale can score highly without measuring a single construct.
References
- AQA. (2021). AS and A-level Psychology (7181, 7182) specification (Version 1.2). Assessment and Qualifications Alliance.
- AQA. (2023). A-level Psychology 7182/1 Paper 1: Introductory topics in psychology. Mark scheme, June 2023. Assessment and Qualifications Alliance.
- Bandura, A., Ross, D., & Ross, S. A. (1961). Transmission of aggression through imitation of aggressive models. Journal of Abnormal and Social Psychology, 63(3), 575–582.
- Brown, W. (1910). Some experimental results in the correlation of mental abilities. British Journal of Psychology, 3(3), 296–322.
- Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46.
- Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297–334.
- Hedge, C., Powell, G., & Sumner, P. (2018). The reliability paradox: Why robust cognitive tasks do not produce reliable individual differences. Behavior Research Methods, 50(3), 1166–1186.
- Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155–163.
- Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174.
- Loftus, E. F., & Palmer, J. C. (1974). Reconstruction of automobile destruction: An example of the interaction between language and memory. Journal of Verbal Learning and Verbal Behavior, 13(5), 585–589.
- McHugh, M. L. (2012). Interrater reliability: The kappa statistic. Biochemia Medica, 22(3), 276–282.
- OCR. (2026). A Level Psychology H567 specification (Version 1.5). Oxford Cambridge and RSA Examinations.
- Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716.
- Pearson Edexcel. (2026). Pearson Edexcel Level 3 Advanced GCE in Psychology (9PS0) specification (Issue 4). Pearson Education.
- Regier, D. A., Narrow, W. E., Clarke, D. E., Kraemer, H. C., Kuramoto, S. J., Kuhl, E. A., & Kupfer, D. J. (2013). DSM-5 field trials in the United States and Canada, Part II: Test-retest reliability of selected categorical diagnoses. American Journal of Psychiatry, 170(1), 59–70.
- Roberts, B. W., & DelVecchio, W. F. (2000). The rank-order consistency of personality traits from childhood to old age: A quantitative review of longitudinal studies. Psychological Bulletin, 126(1), 3–25.
- Spearman, C. (1910). Correlation calculated from faulty data. British Journal of Psychology, 3(3), 271–295.
- Tavakol, M., & Dennick, R. (2011). Making sense of Cronbach's alpha. International Journal of Medical Education, 2, 53–55.
Further Reading and Research
Recommended Articles
- Sampling Methods in Psychology: Techniques and Examples
- Aims and Hypotheses in Psychology: Directional and Non-Directional
- Choosing the Right Statistical Test: A Decision Guide
Suggested Books
- Coolican, H. (2019). Research Methods and Statistics in Psychology (7th ed.). Routledge.
- The standard research methods textbook for A level and undergraduate psychology, with clear chapters on reliability, validity and psychometric testing.
- Field, A. (2024). Discovering Statistics Using IBM SPSS Statistics (6th ed.). Sage.
- A readable, humorous guide to statistics that includes step-by-step reliability analysis and Cronbach's alpha for students moving on to university.
- Streiner, D. L., Norman, G. R., & Cairney, J. (2015). Health Measurement Scales: A Practical Guide to Their Development and Use (5th ed.). Oxford University Press.
- A practical guide to building questionnaires and scales, with detailed chapters on reliability, kappa and the intraclass correlation.
Recommended Websites
- AQA A-level Psychology (7182)
- The official specification, past papers and mark schemes, including the June 2023 question on improving the reliability of a content analysis.
- British Psychological Society
- The professional body for UK psychology, with guidance on research ethics and standards for psychological testing.
- Open Science Framework
- The platform behind the Reproducibility Project, where the data and materials from the Open Science Collaboration's replication attempts are publicly available.
To cite this article please use:
Early Years TV Reliability in Psychology: Types, Assessment and Improvement. Available at: https://www.earlyyears.tv/reliability-psychology-types-assessment-improvement/ (Accessed: 2 October 2026).

