What is Criterion Validity?

Criterion validity is the extent to which a survey or questionnaire predicts or correlates with an external, real-world outcome (a “criterion”). It shows whether the measure actually reflects the thing you care about.

Criterion validity asks: do scores from this survey match something meaningful outside the survey? There are two common types: concurrent validity (the survey matches an outcome measured at the same time, like a clinical diagnosis) and predictive validity (the survey predicts a future outcome, like later service use). Practically, researchers compare survey scores with a trusted external standard using correlations, receiver operating characteristic (ROC) curves, or regression models to see how well the survey identifies or predicts the criterion.

Usage example

A local council creates a short housing-stress questionnaire. To check criterion validity they compare respondents’ questionnaire scores to an independent caseworker assessment done the same week (concurrent validity). They also track whether high scores predict emergency rehousing requests in the next six months (predictive validity).

Practical application

Criterion validity matters because it shows whether a survey produces results that are useful for decisions and action. If a questionnaire has strong criterion validity you can trust its scores to identify people who need follow-up, to evaluate programs, or to allocate resources. For multilingual contexts, maintaining criterion validity means translations must preserve the relationship between questions and the real-world outcome — so you can compare and combine responses across languages with confidence. Practical steps include choosing a clear external criterion, collecting paired data, calculating the appropriate statistics, and rechecking validity after translation edits or major wording changes.

FAQ

How is criterion validity different from construct validity?

Construct validity asks whether a survey measures the intended theoretical concept (is this really measuring “wellbeing” as defined?), while criterion validity asks whether the survey relates to a concrete external outcome (does the score match a clinical diagnosis or predict future behaviour?). Both matter, but criterion validity ties the measure to practical, real-world usefulness.

How do you test criterion validity in practice?

Pick a trusted external standard (an administrative record, professional assessment, or future event), collect survey scores and criterion data from the same sample, and use statistics like correlations, ROC curves, sensitivity/specificity, or regression to quantify the relationship. For predictive validity, you follow participants over time and test whether scores forecast the outcome.

Can translations or language differences affect criterion validity?

Yes. If a translated question shifts meaning or emphasis, its relationship with the external criterion can weaken or change. In multilingual projects, test criterion validity for each language when possible, and use participant feedback and admin review to refine wording. One-survey multilingual platforms that let you update translations and re-evaluate responses make this process easier and more reliable.

What if my sample is small — can I still assess criterion validity?

Small samples limit statistical power and precision. You can still do an initial check (look for strong correlations or clear classification patterns), but treat findings as tentative and plan larger follow-up tests. Where possible, combine data over time or across sites while checking consistency across language and demographic subgroups.