What is Item Response Theory?

Item Response Theory (IRT) is a family of statistical models that link a person's unobserved trait (for example satisfaction, ability or attitude) to the probability they give a particular answer to a survey item. IRT describes each question by parameters such as difficulty and discrimination so you can build and compare reliable scales.

Item Response Theory models how individual survey items behave across respondents by estimating the relationship between a latent trait (the thing you want to measure) and the likelihood of specific responses. Rather than treating every question as equally informative, IRT assigns parameters to items: a difficulty or location parameter (where on the trait scale the item is targeted), a discrimination parameter (how well the item distinguishes between people with different trait levels) and, in some models, a guessing or lower-bound parameter. There are different IRT models for binary items (correct/incorrect), graded responses (Likert scales) and nominal choices. IRT also enables checking whether items behave differently across groups — a concept known as differential item functioning (DIF).

Usage example

A local council runs a resident satisfaction survey in multiple languages. Using IRT on the Likert items, analysts find one question that has low discrimination and shows DIF for respondents answering in Romanian. That flags a possible translation or cultural issue; the team reviews and rewords the item and then re-runs the analysis to confirm improved performance.

Practical application

Why it matters: IRT gives you a principled way to build shorter, more reliable surveys and to compare results across groups and languages. Practical uses include:
- Selecting the most informative items so surveys stay short while preserving measurement precision (useful for busy respondents).
- Detecting items that function differently across language versions or demographic groups (DIF), which helps surface translation errors or culturally inappropriate wording.
- Equating scales so scores remain comparable even when different respondents answer different sets of items.
- Supporting adaptive questionnaires that present questions tailored to each respondent's likely trait level.
For Hearo and other multilingual survey platforms, IRT is particularly valuable because it helps distinguish whether differences in responses reflect real differences in the population or artefacts of wording or translation — making the feedback loop for improving translations more data-driven.

FAQ

Is IRT the same as Classical Test Theory (CTT)?

No. CTT treats each item as contributing equally and focuses on total scores and overall reliability. IRT models items individually and estimates parameters that describe how each item functions across the trait continuum. IRT offers item-level diagnostics, allows for varying information across the scale, and is better suited for item selection and adaptive testing.

How can IRT help with multilingual surveys?

IRT can identify items that show differential item functioning (DIF) between language groups — a sign that wording or translation may be changing how people interpret the question. By flagging problematic items, teams can review translations, collect participant feedback, and improve wording to make results more comparable across languages.

Do I need a statistician or large samples to use IRT?

IRT is more complex than simple score summaries and benefits from statistical expertise, but many tools and packages make it accessible. Sample-size needs vary by model and goals: simple models (like the Rasch/1PL) can be useful with a few hundred responses, whereas stable estimates for more complex models or DIF testing often require several hundred to a few thousand responses per group.

Can IRT be used with Likert-style questions and open-text responses?

Yes for Likert-style (graded) items — there are specific graded-response models in IRT. Open-text responses are qualitative and not modelled directly by IRT; however, text can be coded into categories or sentiment scores and then analysed with IRT or other quantitative methods to assess item functioning.