Measurement, Scales & Psychometrics
Measurement, scales and psychometrics cover methods for designing, scoring and evaluating survey questions to measure attitudes, behaviours and constructs.
This category includes terms such as Likert scale, reliability, validity, Cronbach's alpha, factor analysis and item response theory, plus practical concepts like scale development, scoring, measurement error and cross‑language equivalence for multilingual surveys.
Balanced Scale
A balanced scale is a survey response scale arranged symmetrically with equal positive and negative options (and sometimes a neutral midpoint) so responses are not pushed in one direction. It helps reduce measurement bias and makes results easier to interpret and compare.
Classical Test Theory (CTT)
Classical Test Theory (CTT) is a basic framework for understanding measurement error in tests and surveys: each respondent’s observed score equals their true score plus random error. It focuses on estimating reliability (consistency) of scales and identifying weak questions.
Composite Reliability
Composite reliability is a statistic that estimates how consistently a group of survey items measures the same underlying construct. It is considered a more accurate measure of internal consistency than Cronbach’s alpha when items have different loadings.
Confirmatory Factor Analysis (CFA)
Confirmatory Factor Analysis (CFA) is a statistical method for testing whether a set of survey questions measures the theoretical constructs you expect. It checks if responses fit a predefined structure of underlying factors (latent variables).
Construct Validity
Construct validity is the degree to which a survey or test actually measures the theoretical idea (the “construct”) it claims to measure, such as trust, wellbeing or service access. Strong construct validity means the questions and responses reflect that underlying concept, not something else.
Content Validity
Content validity is the extent to which a survey’s questions fully and appropriately cover the topic or concept you intend to measure. It ensures the instrument asks about all important aspects of that concept, in language participants can understand.
Convergent Validity
Convergent validity shows that a measure relates to other measures it should theoretically be related to. It helps confirm that a survey question or scale is actually capturing the concept you intend to measure.
Criterion Validity
Criterion validity is the extent to which a survey or questionnaire predicts or correlates with an external, real-world outcome (a “criterion”). It shows whether the measure actually reflects the thing you care about.
Cronbach's Alpha
Cronbach's alpha is a statistic that estimates the internal consistency (reliability) of a group of survey items intended to measure the same underlying concept. Values range from 0 to 1, with higher values indicating items that behave more consistently as a set.
Differential Item Functioning (DIF)
Differential Item Functioning (DIF) occurs when a survey question (an item) performs differently for different groups of respondents who have the same underlying trait being measured. In other words, people with equal levels of whatever you're trying to measure respond differently because of group membership (language, culture, age, etc.), not because they actually differ on the trait.
Discriminant Validity
Discriminant validity shows that a survey or test measures a distinct concept rather than something else. It tells you that two different scales are not just re‑measuring the same thing.
Exploratory Factor Analysis (EFA)
Exploratory Factor Analysis (EFA) is a statistical technique used to discover how survey questions group together into underlying themes or ‘factors’. It helps you see whether items measure the same concept without forcing a predefined structure.
Guttman Scale
A Guttman scale (or cumulative scale) is a set of ordered survey items where agreeing with a stronger item implies agreement with all weaker items — allowing a simple, single-number measure of where a respondent sits on a clear, one-dimensional progression.
Interval Scale
An interval scale is a type of measurement where values are ordered and the gaps between values are equal, but there is no meaningful zero point. It lets you compare differences (add/subtract) but not form true ratios (you can say one score is 5 points higher, not 'twice as much').
Inter‑item Correlation
Inter-item correlation measures how closely responses to two survey items move together — it shows whether items intended to measure the same idea are related. It helps diagnose whether items belong together, are redundant, or aren’t measuring the same thing.
Inter‑rater Reliability
Inter-rater reliability (IRR) measures how consistently different people (raters or coders) classify or score the same items. It tells you whether your coding or rating process produces dependable, reproducible results.
Item Characteristic Curve (ICC)
An Item Characteristic Curve (ICC) is a graph that shows how likely people with different levels of an underlying trait (like ability, confidence or agreement) are to give a particular response to one survey question. ICCs help you see where an item is most informative and how well it distinguishes between respondents.
Item Difficulty
Item difficulty is a measure of how easy or hard a survey or test question is for respondents. In tests it’s usually the proportion who answer correctly; in attitude surveys it’s the proportion who endorse (agree with) an item.
Item Discrimination
Item discrimination is a statistic that shows how well a single question (item) separates respondents who score high on the overall questionnaire from those who score low. A high-discrimination item helps the survey measure the intended trait or opinion reliably.
Item Guessing Parameter
The item guessing parameter (often called c) is a value in some item response models that represents the chance a respondent with very low ability will choose a correct answer by guessing. It is used mainly for multiple‑choice or forced‑choice test items to account for lucky guesses.
Item Response Theory (IRT)
Item Response Theory (IRT) is a set of statistical models that describe how a respondent's answer to a survey question relates to their level on an underlying trait (for example ability, attitude or need). IRT helps you evaluate each question’s properties (like how well it discriminates between people) and score respondents on a common scale.
Item–Total Correlation
Item–total correlation measures how well a single survey question (an item) relates to the overall score for the scale it belongs to. It shows whether an item is consistent with the rest of the questions that are intended to measure the same thing.
Likert Scale
A Likert scale is a common survey format that asks respondents to rate their level of agreement, satisfaction or frequency on a ordered set of response options (for example: Strongly disagree → Strongly agree). It’s used to capture attitudes and opinions with simple, comparable responses.
McDonald's Omega
McDonald's Omega is a reliability coefficient that estimates how well a set of survey items measures a single underlying concept. It is generally preferred to Cronbach's alpha when items differ in how strongly they relate to that concept.
Measurement Error
Measurement error is the difference between the information a survey intends to capture and what is actually recorded. It arises when respondents, the question wording, translation, or data collection process distort responses.
Nominal Scale
A nominal scale sorts responses into named categories that have no numeric value or natural order. Examples include language, city, or favourite colour.
Ordinal Scale
An ordinal scale records responses that have a clear order (for example: low → high) but where the gaps between points aren’t assumed equal. It’s commonly used in surveys—most notably in Likert-style questions—to capture ranked opinions or attitudes.
Predictive Validity
Predictive validity is how well a survey question or scale forecasts a future outcome or behaviour. It shows whether responses today actually relate to what happens later.
Rasch Model
The Rasch model is a simple probabilistic measurement model that places people and survey items on the same scale, estimating a respondent's level of a trait (ability, attitude, or risk) and each question's difficulty or endorsement level. It turns ordinal responses (like Likert scales) into interval measures and checks whether items behave consistently.
Ratio Scale
A ratio scale is a numeric measurement scale with equal intervals and a true zero point, so both differences and ratios are meaningful. Examples include weight, time, income and counts.
Reliability
Reliability is how consistently a survey or scale measures the same thing—across items, over time, or between people who score responses. High reliability means results are repeatable and not mainly driven by random error.
Scale Anchors
Scale anchors are the words or labels assigned to points on a rating scale (for example, ‘Strongly disagree’ to ‘Strongly agree’). They tell respondents what each numeric or position value means and shape how answers are interpreted.
Scale Calibration
Scale calibration is the process of making sure a survey scale (for example a 1–5 satisfaction rating) measures the same thing across different groups, languages or settings so scores can be compared fairly. It combines questionnaire design, testing and statistical adjustment to remove systematic differences that aren’t about the underlying opinion or experience.
Semantic Differential Scale
A semantic differential scale measures people’s attitudes by asking them to rate a concept between two opposite adjectives (for example, "helpful — unhelpful") on a numeric scale. It shows both the direction (positive vs negative) and strength of feeling across several bipolar dimensions.
Standard Error of Measurement (SEM)
The Standard Error of Measurement (SEM) estimates how much an individual's observed score on a questionnaire or test is likely to differ from their 'true' score because of measurement error. It shows the typical amount of noise in a score given the test's reliability.
Stapel Scale
The Stapel scale is a rating scale that asks respondents to evaluate a single adjective or attribute by selecting a numeric value (often ranging from negative to positive). It presents one word per item and a row of numbers for respondents to indicate how strongly the adjective applies.
Test‑Retest Reliability
Test-retest reliability measures how consistently a survey or questionnaire produces the same results when the same people answer it more than once under the same conditions. It checks whether responses are stable over time when no real change has occurred.
Thurstone Scale
A Thurstone scale is an attitude-measurement method that uses a set of pre-rated statements and numeric weights to produce an interval-level score for a respondent’s opinion. It relies on a panel of judges to assign values to statements so individual responses can be averaged into a single attitude score.
Validity
Validity is how well a survey or question measures what it is intended to measure. A valid survey produces answers you can trust to represent the concept you care about.
Visual Analog Scale (VAS)
A Visual Analog Scale (VAS) is a continuous response scale — usually a horizontal line or slider with labelled endpoints — that lets respondents mark a point to express the intensity of a feeling or perception (e.g., pain, satisfaction). It captures fine-grained differences that fixed-choice scales can miss.
No glossary terms yet