What is Re-identification Risk?
Re-identification risk is the chance that someone who provided ‘anonymised’ survey data can be identified by combining that data with other information. Even when direct names are removed, small details or combinations of answers can reveal who responded.
Re-identification risk refers to how likely it is that an individual can be identified from data that was supposed to be anonymised. In survey and feedback data this can happen when seemingly harmless details (age, postcode, rare answers, timestamps, or free-text comments) are combined with other public or private datasets to single out a person. Non-experts should know that anonymisation is not a binary state — it is a spectrum of risk — and that context (who has access to the data, what other datasets exist, and how detailed the responses are) determines how safe the data really is.
Usage example
A council runs a resident survey and removes names, but keeps full postcodes, precise ages and verbatim comments. A local journalist cross-checks those postcodes with a social media post and recognises a household. That combination turns anonymised responses back into identifiable records — an instance of re-identification.
Practical application
Understanding re-identification risk matters because respondents can be harmed if sensitive information is linked back to them: privacy breaches, discrimination, or loss of trust. For Hearo users, it affects what you collect, how you store and share responses, and how you report results. Practical steps include collecting only necessary data, aggregating or generalising identifiers (e.g. use broader age bands or partial postcodes), limiting access to raw responses, suppressing or merging small groups in published outputs, and carrying out simple risk assessments before sharing data externally. In multilingual contexts, free-text answers in other languages still carry re-identification risk, so policies for reviewing, redacting and translating open responses remain important.
FAQ
How can I reduce re-identification risk when running multilingual surveys?
Collect the minimum personal detail you need; prefer broader categories (age bands, neighbourhoods) over exact values; limit free-text fields or review and redact identifying details before sharing; suppress or group results for small demographic segments; restrict who can access raw responses and use role-based controls. When translating open-text replies, review them for identifiers before storing or publishing.
Does automatic translation increase re-identification risk?
Translation itself doesn't create new identifiers, but translating open-text responses can reveal wording that makes an individual identifiable (for example, named places or unique life events). Treat translated text the same as original responses: review and redact identifiers, and control access to raw translated content.
What legal obligations do we have around re-identification risk?
Regulations like GDPR require appropriate safeguards when processing personal data and may treat pseudonymised or high-risk data differently. You must justify data collection, apply data minimisation, and implement technical and organisational measures to reduce risk. Follow local data-protection guidance and consider a privacy impact assessment for higher-risk projects.
What should we do if we discover re-identification has happened?
Act quickly: contain further access to the dataset, assess the scope and harm, notify any required authorities and affected individuals per legal requirements, and review mitigation steps (better anonymisation, stricter access controls). Use the incident to update processes so similar breaches are less likely in future.