What is Test‑Retest Reliability?
Test-retest reliability measures how consistently a survey or questionnaire produces the same results when the same people answer it more than once under the same conditions. It checks whether responses are stable over time when no real change has occurred.
Test-retest reliability is a way of checking whether a measurement tool (like a survey question, scale or questionnaire) produces consistent results when given to the same group of people at two different times. If participants’ opinions, abilities or circumstances haven’t changed, a reliable measure will yield similar answers on both occasions. Practically, researchers collect responses at time 1 and again at time 2, then calculate a similarity statistic (for example Pearson’s correlation or intraclass correlation for continuous scores, Cohen’s kappa for categorical items). High agreement suggests the measure is stable; low agreement suggests the item may be ambiguous, sensitive to momentary factors, or poorly translated. Key caveats: the time gap matters (too short and people remember answers; too long and real change can occur), and some questions (e.g., mood or recent events) are naturally less stable.
Usage example
A school runs a parent-engagement survey and repeats the same questionnaire with the same parents after two weeks. The research team calculates the correlation between total scores at both times and finds r = 0.82. They conclude the survey has good test-retest reliability — parents’ responses are stable when their circumstances haven’t changed.
Practical application
Why it matters: Test-retest reliability tells you whether a question or scale is measuring a stable construct rather than noise. In Hearo’s context it helps teams trust that: - scores or ratings reflect real differences between people rather than random fluctuation, - translations retain the original meaning (if a translated version produces similar responses over repeated administration), and - changes observed between survey waves are likely real rather than measurement error. Use test-retest checks when you build a new questionnaire, adapt items into other languages, or when you expect to track the same respondents over time. If reliability is low, review question wording, cultural appropriateness of translations, and the timing of measurement.
FAQ
How long should I wait between tests?
There’s no single rule — choose an interval long enough to avoid simple recall (usually 1–4 weeks for attitude or satisfaction questions) but short enough that you don’t expect real change. The ideal gap depends on the topic: personality can be measured over months, but opinions about a recent event may shift quickly.
What counts as a ‘good’ test-retest score?
Benchmarks vary by field and measure. Correlations or intraclass correlations above ~0.7 are often considered acceptable for group-level work; 0.8–0.9 is strong. For single items, especially categorical ones, lower values are common. Interpret scores alongside sample size, question type and purpose.
Can test-retest check whether translations are accurate?
Partly. If the translated survey yields similar responses on repeat administration, that supports comparability. But test-retest alone can’t prove semantic equivalence; combine it with cognitive testing, participant feedback on wording (a built-in Hearo workflow), and parallel-form checks across languages.
What can make test-retest reliability low, and how do I fix it?
Common causes: vague or double-barrelled questions, sensitive items that genuinely fluctuate, poor translations or culturally inappropriate phrasing, and inappropriate timing. Fixes include simplifying wording, improving translations (review flagged phrases), picking a better interval, piloting with the target community, and using aggregated scales rather than single items when possible.