What is Item Response Theory (IRT)?

Item Response Theory (IRT) is a set of statistical models that describe how a respondent's answer to a survey question relates to their level on an underlying trait (for example ability, attitude or need). IRT helps you evaluate each question’s properties (like how well it discriminates between people) and score respondents on a common scale.

IRT treats each survey question (an item) as having characteristics that determine the probability a person with a particular level of the latent trait will choose a given response. Common item parameters are difficulty (where on the trait scale the item is most informative), discrimination (how sharply the item distinguishes nearby trait levels) and, in some models, a guessing parameter. Unlike simple totals or averages, IRT produces estimated trait scores that account for which items a respondent saw and how informative each item was.

Key assumptions include unidimensionality (items measure a single underlying trait) and local independence (responses to different items are independent once the trait level is controlled for). IRT is widely used in educational tests, health questionnaires and attitude scales, and it supports features such as calibrated item banks and shorter adaptive questionnaires.

Usage example

If a school uses a parental engagement questionnaire, IRT can identify which questions best separate engaged from less-engaged parents. The team can then score parents on a consistent engagement scale even if some respondents answered slightly different item sets or languages.

Practical application

IRT matters because it improves measurement quality and makes multilingual or multi-version surveys more comparable. Practical benefits include:
- Selecting the most informative items so surveys can be shorter without losing precision.
- Producing trait scores that are comparable across respondents even if they saw different items or language versions.
- Detecting items that behave differently for different groups (differential item functioning, or DIF), which can flag translation problems or cultural bias.
- Building calibrated item banks that support adaptive or targeted questionnaires.

For organisations using Hearo, IRT can help validate that translated questions measure the same thing across languages, focus limited translation resources on problematic items, and ensure responses from different communities are comparable and fair.

FAQ

How is IRT different from simply summing questionnaire scores?

Summed scores treat every item as equally informative. IRT models how informative each item is and where on the trait scale it works best, producing scores that reflect both responses and item properties. This yields more precise and comparable measurement, especially when respondents see different items or when items vary in difficulty.

Do I need a large sample to use IRT?

IRT generally needs larger samples than basic summary statistics because it estimates item parameters as well as respondent traits. The required size depends on the model complexity and number of items; simple models and short scales can be useful with a few hundred responses, while more complex calibrations benefit from larger samples.

Can I use IRT to check translations and fairness across languages?

Yes. By comparing item parameters across language groups and testing for differential item functioning (DIF), you can identify items that perform differently and may need rewording or cultural adaptation. That makes IRT a practical tool for improving multilingual survey quality.