What is Item Response Theory (IRT)?
Item Response Theory (IRT) is a family of statistical models that describe how a person’s answer to a survey or test question relates to an underlying trait (like ability, opinion strength or need). IRT models each question by parameters such as difficulty and discrimination to produce fairer scoring and better question diagnostics.
IRT treats each survey question (an item) as having properties that influence the probability a respondent will choose a particular answer, and it treats respondents as having a level on an underlying trait (for example: literacy, agreement, symptom severity). Common IRT parameters are: difficulty (how much of the trait is needed to endorse an item), discrimination (how well the item distinguishes between nearby trait levels), and sometimes guessing (for right/wrong items). Models exist for yes/no or multiple-choice items and for graded responses such as Likert scales. Compared with simple total scores, IRT gives item-level detail, allows comparable scoring even when people answer different sets of items, and helps spot questions that don’t behave consistently across groups.
Usage example
A public health team running a wellbeing survey used IRT to see which questions best separated respondents with low, medium and high wellbeing. They removed items with low discrimination and used the remaining items to generate a single wellbeing score that’s comparable even when some respondents skipped questions. Later, they used the same IRT analysis to check whether translated versions of a question were working differently in different language groups.
Practical application
IRT matters because it gives actionable insight into question quality and fairness. For multilingual surveys, IRT can: identify questions that are poor at measuring the intended concept; detect differential item functioning (DIF) where a translated question performs differently for one language group (a sign of mistranslation or cultural mismatch); enable shorter surveys without losing measurement precision; and produce scores that are comparable across respondents who answered different item sets. In practice this helps teams improve translations, prioritize which items to revise, and trust comparisons between communities that answered in different languages.
FAQ
How is IRT different from adding up item scores (classical test theory)?
Classical test theory uses raw totals and treats all items as equally informative. IRT models each item separately, estimating how informative it is and where on the trait it works best. That gives more precise measurement, allows comparing people who answered different items, and highlights items that may be biased or poorly worded.
Can I use IRT with Likert-style survey questions?
Yes. There are IRT models designed for graded responses (e.g., 5‑point agreement scales). Those models estimate parameters suitable for ordered responses rather than simple right/wrong answers.
Do I need a lot of responses to use IRT?
IRT typically needs more data than simple item analysis. For basic one-dimensional models, hundreds of responses can be enough; more complex models or subgroup analyses (for DIF across languages) need larger samples. Exact needs depend on the model and number of items.
Can IRT help find translation problems in multilingual surveys?
Yes. By comparing item parameters across language groups, IRT can flag items that behave differently (differential item functioning). That’s a strong signal to review the translation or cultural fit of a question.