Validity in Psychology: Types, Assessment and Improvement

Types of validity in psychology table: internal, ecological, population, temporal, face and concurrent, and how to check each

In 1973, eight people with no mental illness claimed to hear voices and were admitted to 12 psychiatric hospitals. Not one was detected by staff, which turned validity in psychology from a textbook term into a real-world problem.

Key Takeaways

  • What is validity? Validity is accuracy: whether a study or measure really tests what it claims to, and whether its findings hold outside the study itself.
  • Which types do I need? Internal and external validity are the two big families. AQA names four types: face, concurrent, ecological and temporal. OCR and Edexcel add population, construct, criterion and predictive validity.
  • How is it assessed and improved? Face validity is judged by eye or by experts, concurrent validity by correlating with an established measure. Validity is improved with control groups, blinding, standardised procedures, anonymity, lie scales and triangulation.

Validity is the evaluation point that turns up in almost every psychology essay. Whenever you write that a study “lacks ecological validity” or “cannot be generalised”, you are making a validity argument. The AQA specification asks for “types of validity across all methods of investigation”, including how validity is assessed and improved, so you need to be able to apply it to experiments, observations, questionnaires, interviews and case studies alike (AQA, 2015).

Validity is the partner of reliability. Reliability asks whether a result is consistent; validity asks whether it is true. A study can repeat perfectly and still be measuring the wrong thing, or measuring the right thing in a situation so artificial that it tells us little about everyday life.

This guide explains what validity means, the difference between internal and external validity, every type named on the A level specifications, how each one is assessed and improved, and how to use validity in exam answers. It draws on real studies throughout, from Loftus and Palmer’s car crash films to the Strange Situation, and finishes with a quiz so you can check you can tell the types apart.

What Is Validity in Psychology?

Validity in psychology is the extent to which a study or measure tests what it claims to test, and the extent to which its findings can be generalised beyond the study. In short, validity is about accuracy and truth, whereas reliability is about consistency.

A simple way to picture it is a darts board. If every dart lands in the same spot, the thrower is consistent, which is reliability. If the darts land on the bullseye, the thrower is accurate, which is validity. A cluster of darts tightly grouped in the corner of the board is reliable but not valid: consistent, but consistently wrong.

Two Questions Validity Asks

Validity always comes back to two questions. The first is whether the researcher really measured what they set out to measure: did the change in the dependent variable come from the independent variable, or from something else? That is internal validity. The second is whether the findings apply beyond this particular study: to other settings, other people and other times. That is external validity.

The two terms were set out by the American psychologist Donald Campbell, who separated the question of whether an experimental effect is real from the question of how far it can be generalised (Campbell, 1957). Every other type of validity on the specification sits under one of these two headings, or describes the accuracy of a single measuring tool.

Validity of Measures and Validity of Studies

Validity can describe a measuring tool or a whole study, and it helps to keep the two apart. A measure, such as a questionnaire, a psychometric test or a set of behavioural categories, is valid if it measures the thing it is named after: an anxiety scale should measure anxiety, not general unhappiness or a wish to look good. Face validity and concurrent validity are judgements about measures.

A study is valid if its conclusions are justified. That depends on the measures, but also on the design: whether extraneous variables were controlled, whether participants guessed the aim, and whether the sample, setting and era allow the findings to be generalised. Internal, ecological, population and temporal validity are judgements about studies.

Types of validity in psychology table: internal, ecological, population, temporal, face and concurrent, and how to check each
Internal validity is about what happens inside a study, and ecological, population and temporal validity are all forms of external validity. Face and concurrent validity judge the measuring tool itself.

Types of Validity in Psychology

The main types of validity in psychology are internal validity and external validity, with external validity divided into ecological, population and temporal validity. Measures are judged by face validity and by criterion validity, which includes concurrent and predictive validity. Construct validity is the overarching question of whether a measure captures the idea it is meant to.

TypeQuestion it answersApplies to
Internal validityDid the IV really cause the change in the DV?Studies
External validityDo the findings generalise beyond this study?Studies
Ecological validityDo the findings generalise to other settings, especially everyday life?Studies
Population validityDo the findings generalise to other people?Studies
Temporal validityDo the findings still hold in other time periods?Studies
Face validityDoes the measure look, on the surface, as if it measures what it should?Measures
Concurrent validityDo scores match an established measure taken at the same time?Measures
Predictive validityDo scores predict a relevant outcome in the future?Measures
Construct validityDoes the measure capture the underlying idea, such as intelligence or anxiety?Measures

What Each Exam Board Requires

The three main A level boards list validity differently, so check which list applies to you. AQA names four types, face, concurrent, ecological and temporal validity, together with assessing and improving validity (AQA specification, section 4.2.3). OCR has the longest list, and Edexcel names three.

BoardTypes of validity named in the specification
AQA (7182)Face, concurrent, ecological and temporal validity; assessment of validity; improving validity. Also validity in the diagnosis of schizophrenia.
OCR (H567)Internal, face, construct, concurrent, criterion, external, population and ecological validity.
Edexcel (9PS0)Internal, predictive and ecological validity, alongside objectivity and reliability.

Even if your board does not name internal and external validity, examiners credit them, and you need the ideas behind them to explain the named types. Ecological and temporal validity only make sense once you see them as kinds of external validity.

Internal Validity: Did the IV Cause the Change?

Internal validity is the extent to which a study measures what it set out to measure, so that any change in the dependent variable (DV) can be put down to the independent variable (IV) and nothing else. A study with high internal validity has ruled out the other explanations.

Internal validity matters most in experiments, because the whole point of an experiment is to show cause and effect. If a rival explanation could account for the result, the experiment has not shown what it claims, however large or significant the difference between conditions.

Threats to Internal Validity

Anything other than the IV that could change the DV is a threat to internal validity. These are the ones that come up most often at A level.

  • Extraneous and confounding variables. An extraneous variable is any variable other than the IV that might affect the DV. When it changes systematically with the IV, it becomes a confounding variable, and the result cannot be trusted. Testing one condition in the morning and the other after lunch is a classic example.
  • Participant variables. In an independent groups design, differences between the people in each group, such as age, ability or motivation, may explain the result. Random allocation is the usual defence.
  • Order effects. In a repeated measures design, practice, boredom or fatigue can change performance in the second condition. Counterbalancing spreads these effects across conditions.
  • Demand characteristics. Cues that tell participants what the study is about can change how they behave, so the DV reflects their guess about the aim rather than the IV.
  • Investigator effects. The researcher’s expectations can leak into the study through tone of voice, body language or the way results are recorded.

Several of these threats are controlled by the choice of design. Our guide to experimental designs explains how repeated measures, independent groups and matched pairs each trade one threat for another.

Demand Characteristics: Orne’s Impossible Task

Demand characteristics are the cues in a study that communicate what the researcher expects, and they are one of the most serious threats to internal validity. The term comes from Martin Orne, who described them as the “sum total” of cues that convey the hypothesis to participants, including rumours about the study, the way it was advertised, the experimenter and the setting (Orne, 1962).

Orne showed how far participants will go to be “good subjects”. He tried to design a task so boring and pointless that people would refuse to continue. Participants were given a stack of about 2,000 sheets of random numbers, each needing 224 additions, and told to keep working. One participant was still working when the experimenter gave up five and a half hours later. In a harsher version, after each sheet participants picked up a card telling them to tear the finished sheet into at least 32 pieces and carry on. They still persisted for hours, treating the task as meaningful because it was part of an experiment (Orne, 1962).

The lesson for validity is that participants are not passive. They look for the purpose of a study and often try to help, so a change in behaviour may reflect what they think the researcher wants rather than the effect of the IV. Social desirability bias, where people give answers that make them look good, is a close relative in questionnaires and interviews.

A Real Example: Loftus and Palmer (1974)

Loftus and Palmer’s car crash study is a good example of high internal validity. In their first experiment, 45 students watched seven films of traffic accidents and were asked how fast the cars were going when they “smashed”, “collided”, “bumped”, “hit” or “contacted” each other. Only the verb changed between conditions, and the mean speed estimates ranged from 40.8 mph for “smashed” to 31.8 mph for “contacted” (Loftus and Palmer, 1974).

Because everything else was held constant, the difference in estimates can be put down to the wording of the question. A second experiment with 150 students found that, a week later, 16 of the 50 people asked the “smashed” question reported seeing broken glass, against 7 of 50 in the “hit” condition and 6 of 50 controls, although the film showed none (Loftus and Palmer, 1974).

The same study is often criticised for low ecological validity. Watching a film clip in a classroom, knowing you are in a study, is not the same as witnessing a real crash, where fear and surprise may change what you notice and remember. Our full guide to Loftus and Palmer’s car crash experiment covers both sides. This is the most common pattern in psychology: tight control buys internal validity at the cost of external validity.

External Validity: Can the Findings Be Generalised?

External validity is the extent to which the findings of a study can be generalised beyond the study itself: to other settings, other people and other time periods. Each of those three directions has its own name, ecological validity, population validity and temporal validity.

External validity does not mean “done in the real world”. A field experiment can still have poor external validity if the sample is narrow, and a laboratory experiment can generalise well if the process it studies works the same way everywhere. The question is always whether the particular features of this study limit how far its conclusions travel.

How Well Do Laboratory Findings Generalise?

Laboratory findings generalise to the field more often than critics assume, but not evenly across psychology. Mitchell (2012) compared effect sizes from laboratory and field studies of the same topics using 217 lab-field comparisons from 82 meta-analyses. Industrial-organisational psychology, the study of people at work, had laboratory results that best predicted field results. Social psychology laboratory effects most often reversed direction in the field, and large laboratory effects were more reliably replicated than medium or small ones (Mitchell, 2012).

That gives you a sharper evaluation point than “lab studies lack external validity”. Whether a laboratory finding generalises depends on the topic and on the size of the effect, and a small effect found under tight control is the one most at risk of disappearing outside the laboratory.

Population Validity

Population validity is the extent to which findings can be generalised from the sample to other groups of people. It depends on how the sample was chosen: a small volunteer sample of psychology undergraduates cannot safely represent children, older adults or people from other cultures.

Population validity is a real weakness of psychology as a whole, not just of individual studies. An analysis of the top journals in six areas of psychology from 2003 to 2007 found that 96% of participants came from Western industrialised countries, which hold only 12% of the world’s population. In the leading social psychology journal, 67% of American samples were made up solely of psychology undergraduates (Henrich et al., 2010, drawing on Arnett, 2008).

Henrich and colleagues called these samples WEIRD: Western, Educated, Industrialised, Rich and Democratic. Their review found that WEIRD participants are often outliers rather than typical humans, on topics from visual perception to fairness and moral reasoning (Henrich et al., 2010). This links population validity directly to gender and cultural bias in psychology, and to how a sample is chosen in the first place, which is covered in our guide to sampling methods.

Temporal Validity

Temporal validity is the extent to which findings from one period of history still apply today, or in other periods. Studies of behaviour shaped by social norms are the most exposed, because norms about obedience, gender, conformity and mental health change over decades.

The best-known example is conformity. In the 1950s, Asch’s line study found that participants often went along with a unanimous but obviously wrong majority. When Perrin and Spencer replicated the procedure with British students around 1980, compliance appeared on only one of 396 critical trials. They argued that the Asch effect reflected the United States of the 1950s (Perrin and Spencer, 1981).

Their findings also show why temporal validity is not the whole story. In the same paper, levels of compliance similar to Asch’s were found among young people on probation when the majority and the experimenter were probation officers (Perrin and Spencer, 1981). Conformity had not vanished; it depended on the cost of standing out. A good exam answer uses both halves of that finding.

Temporal validity also applies to definitions and diagnoses. AQA’s mark scheme credits “social norms change over time (lack of temporal validity)” as an evaluation of deviation from social norms as a definition of abnormality (AQA, 2021). Our guide to defining abnormality shows how this plays out.

Ecological Validity: Does It Hold in Real Life?

Ecological validity is the extent to which findings can be generalised to other settings, especially everyday life. It is a type of external validity and the one that students use most, usually to criticise laboratory experiments for being artificial.

Ecological validity depends on more than the setting. The task matters too: remembering a list of unrelated words is not like remembering a conversation, and giving electric shocks is not like most everyday acts of obedience. A study can be run in a natural setting and still use an artificial task, or run in a laboratory with a task that closely matches real behaviour.

Mundane Realism

Mundane realism is how closely the tasks and materials in a study resemble things people do in everyday life. Low mundane realism is one reason a study may lack ecological validity, but the two are not the same thing. Mundane realism describes the study; ecological validity describes whether its findings generalise.

Milgram’s obedience study is the standard example. Twenty-six of his 40 participants, 65%, went on to the highest shock level on the generator (Milgram, 1963), but ordering a stranger to shock a learner has little in common with the orders most people receive at work or school. Critics argue that this low mundane realism limits how far the 65% figure tells us about everyday obedience. Our guide to Milgram’s shock experiment sets out the debate.

When Lab and Field Disagree: The Nurse Studies

A useful way to test ecological validity is to compare what people do in a controlled study with what they do in a natural setting. Hofling and colleagues ran a field experiment in a real hospital. A caller claiming to be a doctor phoned nurses and asked them to give a patient twice the maximum dose of an unfamiliar drug, “Astroten”. Of the 22 nurses, 21 were stopped on their way to give the dose (Hofling et al., 1966).

That looks like strong evidence of obedience in a real workplace. But when Rank and Jacobson (1977) repeated the idea with a familiar drug, Valium, and nurses were free to talk to colleagues as they normally would, most nurses did not comply. The title of their paper calls it “a failure to replicate”. The lesson is that a natural setting alone does not guarantee ecological validity: Hofling’s nurses faced an unknown drug and no chance to check with anyone, which was not how the hospital usually worked.

Is a Laboratory Always Low in Ecological Validity?

A laboratory is not automatically low in ecological validity, and the developmental psychologist Urie Bronfenbrenner made this point in 1977. He defined ecological validity as the extent to which the environment experienced by participants has the properties the investigator assumes it has (Bronfenbrenner, 1977, as cited in Holleman et al., 2020). On that view, the question is whether the setting suits the research question, not whether it looks like real life.

The Strange Situation shows why. If the question is how a young child behaves when placed in an unfamiliar room and separated from a caregiver, a laboratory room is exactly the right setting. If the question is how the child behaves at home, it is not (Holleman et al., 2020). AQA’s mark scheme accepts “controlled observation lacks ecological validity” as an evaluation of the Strange Situation (AQA, 2021), but the strongest answers explain which conclusions the setting limits, and which it does not.

The term itself has moved a long way from its origins. Egon Brunswik introduced “ecological validity” in the 1940s and 1950s to describe how well a cue in the environment predicts something else, a correlation rather than a verdict on a study. Researchers at Utrecht University argue that today’s “real-world” meaning is so loose that it is often used as a blunt criticism without saying which setting a study should generalise to (Holleman et al., 2020). You do not need Brunswik for the exam, but specifying the setting is exactly what examiners reward.

Face Validity and Concurrent Validity

Face validity and concurrent validity are both ways of judging whether a measure, such as a questionnaire or psychological test, measures what it claims to. Face validity is a surface judgement; concurrent validity is a statistical check against an established measure.

Face Validity

Face validity is whether a measure appears, at first sight, to measure what it is supposed to. A question such as “How often do you feel worried?” has obvious face validity as an item on an anxiety scale. A question about favourite colours does not.

Face validity is assessed by simply looking at the measure, or by asking an expert in the field to check it. It is the weakest form of validity because it rests on judgement, not evidence: a measure can look right and still be flawed. It also cuts both ways. Items that are very obviously about the topic make it easy for participants to guess the aim and give socially desirable answers, which is why some questionnaires hide their true purpose among filler items.

Concurrent Validity

Concurrent validity is whether scores on a new measure match scores on an established, already validated measure of the same thing, when both are taken at about the same time. If the two sets of scores are closely related, the new measure is probably measuring the same thing.

Concurrent validity is assessed with a correlation coefficient. The same group of participants completes both measures, and the two sets of scores are correlated. A strong positive correlation is evidence of concurrent validity. A level teaching materials usually borrow the +0.80 benchmark that AQA mark schemes use for reliability, so a correlation of +0.80 or above is a sensible figure to quote.

A real example comes from early years practice. The Strengths and Difficulties Questionnaire (SDQ), a short behavioural screening questionnaire widely used with children, was checked against the longer, established Rutter questionnaires. Parents and teachers of 403 children from dental and psychiatric clinics completed both, and the scores were highly correlated. The two measures were also equally good at telling psychiatric and dental clinic attenders apart (Goodman, 1997). A later study compared the SDQ with the Child Behavior Checklist for 132 children aged four to seven, with the same result (Goodman and Scott, 1999).

Concurrent validity has one obvious limit: it is only as good as the established measure. If the older questionnaire is itself flawed, agreeing with it proves little.

Criterion, Predictive and Construct Validity

Criterion validity is the umbrella term for checking a measure against an outside standard, or criterion. Concurrent validity is one form, where the criterion is measured at the same time. Predictive validity is the other, where the criterion comes later: an aptitude test has predictive validity if high scorers go on to perform well in the job or course it was designed for.

Construct validity is the broadest question of all: does the measure capture the underlying idea, or construct, it is meant to? Constructs such as intelligence, extraversion or anxiety cannot be observed directly, so there is no single criterion to check against. Cronbach and Meehl (1955) argued that construct validity is built up from a whole network of evidence: the measure should relate to other measures and behaviours in the ways the theory predicts. Personality tests are the classic case, and our article on MBTI test accuracy shows what happens when that evidence is weak.

How Is Validity Assessed?

Validity is assessed in different ways depending on the type. Face validity is judged by looking at the measure or asking an expert. Concurrent and predictive validity are assessed by correlating scores with a criterion. Internal validity is judged by checking the design for uncontrolled variables and cues, and external validity by replicating the study in new settings, samples and time periods.

TypeHow it is assessedWhat counts as good evidence
Face validityThe researcher, or an independent expert, examines the items or procedureExperts agree the items clearly relate to the topic
Concurrent validityParticipants complete the new measure and an established one; scores are correlatedA strong positive correlation, usually +0.80 or above
Predictive validityScores are correlated with a later outcomeA strong positive correlation with the outcome
Internal validityThe design is checked for confounding variables, demand characteristics and investigator effectsRival explanations have been ruled out by control
Ecological validityThe study is repeated in a different setting, or compared with field dataSimilar findings in the new setting
Population validityThe study is repeated with a different sampleSimilar findings in different groups
Temporal validityThe study is repeated in a later periodSimilar findings years or decades later

Notice that three of these methods are forms of replication. That is why validity is so closely tied to replicability: a finding that only appears in one setting, one sample or one decade has limited external validity, however carefully the original study was run.

How to Improve Validity in Psychology

Validity is improved by removing the rival explanations for a result and by making the study, its measures and its sample closer to what the findings are meant to describe. The best improvement always depends on the method, so tailor your answer to the one in the question.

Improving Validity in Experiments

  • Use a control group or condition. Comparing the experimental condition with a control shows whether the IV caused the change, or whether it would have happened anyway, through practice or the passage of time.
  • Single-blind procedures. Participants are not told the aim or which condition they are in, which reduces demand characteristics.
  • Double-blind procedures. Neither the participants nor the person collecting the data knows who is in which condition, which also removes investigator effects. This is standard in drug trials, where a placebo group is compared with the treatment group.
  • Standardised procedures and instructions. Every participant has the same experience apart from the IV.
  • Random allocation and counterbalancing. These deal with participant variables and order effects.
  • More realistic tasks and settings. Where the aim is to generalise to everyday life, raising mundane realism or running a field experiment can improve ecological validity.

Improving Validity in Questionnaires and Interviews

  • Guarantee anonymity. Participants who know their answers cannot be traced are more likely to be honest, which reduces social desirability bias.
  • Include a lie scale. Items such as “I have never told a lie” catch respondents who are presenting themselves too favourably, and can make it less obvious what is being measured.
  • Check concurrent validity. Compare scores with an established questionnaire and remove or rewrite items that do not fit.
  • Remove problem items. Leading, ambiguous, overly complex or double-barrelled questions should be rewritten or removed, ideally after a pilot study.
  • Train interviewers. Neutral, consistent questioning reduces investigator effects in interviews.

Improving Validity in Observations

  • Covert observation. People who do not know they are being watched behave more naturally, which raises ecological validity. This raises ethical issues about consent.
  • Clear, operationalised behavioural categories. Categories that do not overlap and cover all the behaviour of interest mean observers record what the study claims to measure.
  • Naturalistic settings. Observing behaviour where it normally happens improves ecological validity, at some cost to control.

Improving Validity in Qualitative Research

Qualitative methods, such as unstructured interviews, case studies and thematic analysis, are often said to have high ecological validity because they capture people’s experiences in depth and in their own words. Their weakness is interpretation: the researcher decides what the data mean.

  • Triangulation. Comparing findings from several sources, such as interviews, observations, diaries and records, and checking that they agree.
  • Direct quotations. Reporting participants’ own words lets readers judge the interpretation for themselves.
  • Reflexivity. The researcher acknowledges how their own views might have shaped the analysis.
  • Checking with participants. Asking participants whether the interpretation fits their experience.

Validity vs Reliability

Validity is about accuracy and reliability is about consistency. A valid measure measures what it claims to; a reliable measure gives the same result each time. The link runs one way: a measure cannot be valid if it is unreliable, but a reliable measure is not necessarily valid.

ValidityReliability
Core questionIs it accurate?Is it consistent?
Main typesInternal, external (ecological, population, temporal), face, concurrentInternal, external (test-retest, inter-observer)
Assessed byExpert judgement, correlation with a criterion, replication in new contextsTest-retest, inter-observer checks, split-half
Improved byControl groups, blinding, anonymity, lie scales, triangulationStandardisation, clear categories, trained observers, pilot studies

Some improvements help one and hurt the other. Moving a study out of the laboratory into a natural setting may raise ecological validity but makes the procedure harder to standardise, so reliability can fall. Our full guide to reliability in psychology covers the consistency side in detail.

Validity in Diagnosis: Rosenhan’s Pseudopatients

Validity in diagnosis is whether a diagnosis identifies a real condition that the person actually has. AQA asks about it directly for schizophrenia, and the most famous study of it is David Rosenhan’s “On being sane in insane places“.

Eight pseudopatients, Rosenhan among them, sought admission to 12 hospitals in five US states. The group included a psychology graduate student, three psychologists, a paediatrician, a psychiatrist, a painter and a housewife. Each said they heard voices saying “empty”, “hollow” and “thud”, then behaved normally once admitted. All but one were admitted with a diagnosis of schizophrenia, and each was discharged with schizophrenia “in remission”. Stays ranged from 7 to 52 days, with an average of 19 (Rosenhan, 1973).

Staff never detected them, but other patients often did: on the first three admissions, 35 of 118 patients voiced suspicions. Rosenhan then told a teaching hospital that pseudopatients would try to gain admission over the next three months. Of 193 patients admitted, 41 were judged with high confidence to be pseudopatients by at least one member of staff, and 19 were suspected by a psychiatrist and one other staff member. In fact, Rosenhan had sent nobody (Rosenhan, 1973).

Rosenhan’s study became a landmark argument that psychiatric diagnosis lacked validity. It now carries a warning of its own. The journalist Susannah Cahalan could trace only two of the other pseudopatients, and one of them had a positive hospital experience that was left out of the published results (Abbott, 2019). The study is still worth knowing, but a strong answer notes that its own validity has been questioned.

Quiz: Spot the Type of Validity

The quickest way to lose marks on validity is to name the wrong type. Work through these eight scenarios and pick the type of validity each one is about.

Spot the Type of Validity

Eight research scenarios. Decide which type of validity each one is about.

Scenario 1 of 8

A new ten-minute reading test is given to 60 children on the same day as a well-established reading assessment, and the two sets of scores are correlated.

How to Answer Validity Exam Questions

Validity questions reward answers that are specific to the study in front of you. Name the type of validity, explain exactly what feature of the study causes the problem, and say what that means for the conclusions.

A real example from AQA Paper 1 asked students to "explain how the validity of the researcher's questionnaire could be improved" for four marks. The mark scheme credited comparing the questionnaire with an existing one and noting differences, removing items that are leading, ambiguous, too complex or double-barrelled, and adding a lie scale so respondents are less aware of what is being tested. It also stated that suggestions about the design of the study were not creditworthy (AQA, 2021).

That last line is the trap. The question was about the questionnaire, so talking about sample size or random allocation earned nothing. A strong answer might read: "The researcher could check concurrent validity by giving participants an established locus of control questionnaire as well, correlating the two sets of scores. Items on the new questionnaire that do not fit the established measure could then be removed or reworded, and a correlation of +0.80 or above would suggest the new questionnaire measures the same thing."

Common Mistakes to Avoid

  • Writing "it lacks ecological validity" and stopping. Say which feature is artificial and which real-life behaviour the findings might not apply to.
  • Treating every lab study as invalid. Control is what gives experiments their internal validity. Weigh the trade-off rather than dismissing the study.
  • Mixing up validity and reliability. "The study was standardised so it can be repeated" is a reliability point.
  • Confusing mundane realism with ecological validity. Mundane realism describes the task; ecological validity describes whether the findings generalise.
  • Using "generalisability" vaguely. Say whether you mean other settings, other people or other times, and use the matching term.

Conclusion

Validity asks whether psychology's findings are true, not just repeatable. Internal validity is about whether a study showed what it claims, and external validity about whether that holds for other settings, people and times. Face and concurrent validity judge the measures themselves. Learn the types, match the improvement to the method, and always name the specific feature of a study that raises or lowers its validity.

Frequently Asked Questions

What is validity in psychology?

Validity in psychology means accuracy. A valid study or measure tests what it claims to test, and a valid finding can be generalised beyond the study. It is usually split into internal validity, which concerns cause and effect within a study, and external validity, which concerns whether results apply to other settings, people and times.

What is the difference between internal and external validity?

Internal validity is about what happens inside a study: whether the independent variable, and not some other factor, produced the result. External validity is about what happens outside it: whether the result generalises to other places, populations and periods. Tight laboratory control tends to raise internal validity while narrowing external validity.

What is ecological validity in psychology?

Ecological validity is a form of external validity concerned with settings. It asks whether findings from a study, often a laboratory experiment, apply to behaviour in everyday situations. Both the environment and the task matter, so an artificial task can lower ecological validity even when the study takes place somewhere natural.

What is temporal validity in psychology?

Temporal validity is whether research findings stay true across different eras. Findings tied to social attitudes, such as conformity or obedience, can date quickly as those attitudes shift. It is tested by repeating an older study years later and comparing the results.

What is population validity?

Population validity is whether findings from one sample apply to other groups of people. It is low when a sample is narrow, for example all university students, all men or all from one culture. Representative sampling and repeating studies with different groups improve it.

What is the difference between face validity and content validity?

Face validity is a quick surface check that a test seems to measure its topic. Content validity goes further, asking whether the items cover every important part of that topic, usually judged systematically by experts. A depression scale that asks only about sadness might have face validity but poor content validity, because it misses sleep, appetite and energy.

What is the difference between concurrent and predictive validity?

Concurrent and predictive validity are both forms of criterion validity; the difference is timing. Concurrent validity compares test scores with a criterion measured at the same time, such as an established test. Predictive validity compares them with a criterion measured later, such as exam grades or job performance.

What is the difference between mundane realism and ecological validity?

Mundane realism is a feature of a study's design: how far its tasks resemble ordinary activities. Ecological validity is a property of its findings: whether they apply in real settings. Low mundane realism is a common reason for low ecological validity, but a study can use unusual tasks and still produce findings that generalise.

How can you improve the validity of a questionnaire in psychology?

  1. Keep responses anonymous so people answer honestly.
  2. Add a lie scale to detect answers given to look good.
  3. Correlate scores with an established questionnaire.
  4. Pilot the questionnaire and rewrite leading or unclear items.

References

  • Abbott, A. (2019). On the troubling trail of psychiatry's pseudopatients stunt. Nature, 574(7780), 622-623.
  • AQA. (2015). AS and A-level Psychology (7181, 7182) specification. Assessment and Qualifications Alliance.
  • AQA. (2021). A-level Psychology 7182/1 Paper 1: Introductory topics in psychology. Mark scheme. Assessment and Qualifications Alliance.
  • Arnett, J. J. (2008). The neglected 95%: Why American psychology needs to become less American. American Psychologist, 63(7), 602-614.
  • Bronfenbrenner, U. (1977). Toward an experimental ecology of human development. American Psychologist, 32(7), 513-531.
  • Campbell, D. T. (1957). Factors relevant to the validity of experiments in social settings. Psychological Bulletin, 54(4), 297-312.
  • Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302.
  • Goodman, R. (1997). The Strengths and Difficulties Questionnaire: A research note. Journal of Child Psychology and Psychiatry, 38(5), 581-586.
  • Goodman, R., & Scott, S. (1999). Comparing the Strengths and Difficulties Questionnaire and the Child Behavior Checklist: Is small beautiful? Journal of Abnormal Child Psychology, 27(1), 17-24.
  • Henrich, J., Heine, S. J., & Norenzayan, A. (2010). The weirdest people in the world? Behavioral and Brain Sciences, 33(2-3), 61-83.
  • Hofling, C. K., Brotzman, E., Dalrymple, S., Graves, N., & Pierce, C. M. (1966). An experimental study in nurse-physician relationships. Journal of Nervous and Mental Disease, 143(2), 171-180.
  • Holleman, G. A., Hooge, I. T. C., Kemner, C., & Hessels, R. S. (2020). The 'real-world approach' and its problems: A critique of the term ecological validity. Frontiers in Psychology, 11, 721.
  • Loftus, E. F., & Palmer, J. C. (1974). Reconstruction of automobile destruction: An example of the interaction between language and memory. Journal of Verbal Learning and Verbal Behavior, 13(5), 585-589.
  • Milgram, S. (1963). Behavioral study of obedience. Journal of Abnormal and Social Psychology, 67(4), 371-378.
  • Mitchell, G. (2012). Revisiting truth or triviality: The external validity of research in the psychological laboratory. Perspectives on Psychological Science, 7(2), 109-117.
  • Orne, M. T. (1962). On the social psychology of the psychological experiment: With particular reference to demand characteristics and their implications. American Psychologist, 17(11), 776-783.
  • Perrin, S., & Spencer, C. (1981). Independence or conformity in the Asch experiment as a reflection of cultural and situational factors. British Journal of Social Psychology, 20(3), 205-209.
  • Rank, S. G., & Jacobson, C. K. (1977). Hospital nurses' compliance with medication overdose orders: A failure to replicate. Journal of Health and Social Behavior, 18(2), 188-193.
  • Rosenhan, D. L. (1973). On being sane in insane places. Science, 179(4070), 250-258.

Further Reading and Research

Recommended Articles

Suggested Books

  • Coolican, H. (2019). Research Methods and Statistics in Psychology (7th ed.). Routledge.
    • The standard research methods textbook for A level and undergraduate psychology, with detailed chapters on every type of validity and on threats to experimental validity.
  • Cahalan, S. (2019). The Great Pretender: The Undercover Mission That Changed Our Understanding of Madness. Grand Central Publishing.
    • A journalist's investigation of Rosenhan's pseudopatient study, and a gripping case study in questioning the validity of a famous piece of research.
  • Henrich, J. (2020). The WEIRDest People in the World: How the West Became Psychologically Peculiar and Particularly Prosperous. Farrar, Straus and Giroux.
    • The book-length follow-up to the WEIRD paper, explaining why Western participants differ psychologically and what that means for population validity.

Recommended Websites

  • AQA A-level Psychology (7182)
    • The official specification, past papers and mark schemes, including the Paper 1 question on improving the validity of a questionnaire.
  • British Psychological Society
    • The professional body for UK psychology, with the Code of Human Research Ethics and guidance on standards for psychological tests.
  • Frontiers in Psychology
    • An open-access journal where you can read the full Holleman and colleagues critique of ecological validity, and other methods papers, free of charge.

Kathy Brodie

Kathy Brodie is an Early Years Professional, Trainer and Author of multiple books on Early Years Education and Child Development. She is the founder of Early Years TV and the Early Years Summit.

Kathy's Author Profile
Kathy Brodie

To cite this article please use:

Early Years TV Validity in Psychology: Types, Assessment and Improvement. Available at: https://www.earlyyears.tv/validity-psychology-types-assessment-improvement/ (Accessed: 3 October 2026).