What is Item response theory (IRT)?
Item Response Theory (IRT) is a family of statistical models that describe how individual survey or test items relate to an underlying trait (like ability, attitude or satisfaction). IRT models estimate item properties (difficulty, discrimination, sometimes guessing) and place respondents on a common scale based on their pattern of answers.
IRT treats each question (item) as a probabilistic function of an unobserved characteristic — often called a latent trait or ability (θ). Instead of simply summing scores, IRT models the probability that a person with a given θ will endorse or correctly answer an item, based on item parameters: difficulty (how much of the trait is needed), discrimination (how well the item differentiates nearby trait levels), and sometimes a guessing parameter (chance of a correct response irrespective of trait). Models range from simple (Rasch or 1-parameter) to more complex (2-parameter, 3-parameter) and include variants for ordered response scales (graded response, partial credit). IRT is widely used for tests, surveys with ordered response options, and any measurement where you want more precise scaling or to compare items and respondents on a common latent scale.
Usage example
A school running a parent engagement survey uses IRT to check which questions best distinguish between families who are highly engaged and those less engaged. The team removes low-discrimination items and shortens the survey while preserving measurement accuracy.
Practical application
IRT matters because it produces more accurate, interpretable measurement than simple total scores. Practical benefits include: enabling shorter surveys without losing precision, creating comparable scores across different item sets or language versions, detecting items that behave differently for subgroups (differential item functioning) — useful for spotting translation or cultural bias — and powering adaptive questionnaires that ask fewer questions tailored to each respondent. For multilingual and community engagement work, IRT helps ensure that translated items measure the same thing across languages and that reported differences reflect real variation rather than measurement artifacts. Note: IRT requires sufficient sample sizes and some technical setup, and assumes items measure a common underlying trait (unidimensionality).
FAQ
How is IRT different from adding up item scores (classical test theory)?
Classical scoring treats each item equally and produces a raw total; IRT models each item’s properties and estimates a respondent’s position on a latent scale. That yields more precise scores, allows comparison across different item sets, and identifies items that aren’t functioning well.
Can I use IRT with Likert-style survey questions?
Yes. Polytomous IRT models (like the graded response or partial credit models) are designed for ordered response options such as Likert scales and estimate how category thresholds relate to the latent trait.
Will IRT help with multilingual surveys?
Yes. IRT can detect differential item functioning (DIF), which flags items that perform differently across language groups — a strong signal that a translation or cultural issue may exist. It also lets you place respondents on a common scale even when they answer slightly different item sets.
Is IRT difficult to implement?
Implementing IRT requires statistical software and enough response data, plus checks for assumptions like unidimensionality and local independence. Many tools and consultants can help; for practical use, teams often apply IRT for analysis and item calibration rather than day-to-day survey creation.