What is Item Difficulty?

Item difficulty is a measure of how easy or hard a survey or test question is for respondents. In tests it’s usually the proportion who answer correctly; in attitude surveys it’s the proportion who endorse (agree with) an item.

Item difficulty quantifies how many people give a particular response to a single question. In classical testing (right/wrong items) it’s calculated as the proportion of respondents who answer the item correctly (often called the p-value). For non‑test questions (Likert, agreement, behavior frequency) the same idea is used to describe how commonly an item is endorsed — for example, how many people agree with a statement or report doing a behavior. Very high or very low difficulty values signal ceiling or floor effects (almost everyone gives the same answer), while intermediate values indicate more spread that helps discriminate between respondents.

Usage example

A knowledge item answered correctly by 80 out of 100 respondents has an item difficulty of 0.80 (relatively easy). A survey statement where only 15 out of 100 respondents agree has a difficulty/endorsement of 0.15 (hard to endorse). If an item is unexpectedly easy in one language and hard in another, that flags a possible translation or cultural issue.

Practical application

Why it matters: Item difficulty helps you spot questions that give little useful information (too easy/hard) and balance a scale or test for better measurement. For Hearo’s multilingual surveys, tracking difficulty across languages is especially important — differences can reveal mistranslation, cultural mismatch, or sampling differences. Use item difficulty to: identify items that need rewriting or clearer translation; avoid floor/ceiling effects that reduce analytic power; choose items that discriminate well for your target population; and run follow-ups (qualitative checks or differential item functioning analysis) when items behave differently across language groups.

FAQ

How is item difficulty calculated?

For right/wrong items it’s the proportion correct: number correct divided by total responses (a value between 0 and 1). For agreement or frequency items you can compute the proportion endorsing the target response (e.g., % who agree or % who say 'often').

What is the ideal item difficulty?

There’s no single ideal for all uses. For tests that need to discriminate, mid-range values (around 0.4–0.6) often provide the most information. For attitude items you want variability without everyone clustering at one end; avoid items with very low (<0.2) or very high (>0.8) endorsement unless those extremes are purposeful.

If an item has different difficulty across languages, does that mean the translation is wrong?

Not always, but it’s a clear sign to investigate. Differences can result from translation issues, cultural interpretation, sample differences, or true variation in experiences. Use participant feedback, review translations, and, when possible, run differential item functioning checks to determine the cause.

Can I fix items that are too easy or too hard?

Yes. Options include rewriting the question for clarity, changing response options, adding follow-up items, or splitting a compound item. In multilingual surveys, also review and revise translations and use participant flagging to collect suggestions for better wording.