What is Post-Stratification?

Post-stratification is a weighting technique applied after data collection to make a survey sample better match known characteristics of the target population (for example age, gender, region or language). It adjusts each response so summary results reflect the population proportions.

Post-stratification is used when a survey’s respondents don’t exactly match the target population on one or more key characteristics. You define categories (strata) for those characteristics — such as age groups, gender, or preferred language — compare the proportion of each category in your sample to the known population proportions, and assign a weight to each respondent so that weighted totals match the population benchmarks. Practically, a simple weight = (population proportion for the stratum) / (sample proportion for the stratum). Variants include raking (iterative proportional fitting) when you need to adjust simultaneously for several variables.

Usage example

A city runs an online consultation and finds 20% of respondents are aged 65+, but the city population is 35% aged 65+. Using post-stratification, the responses from people aged 65+ are given a larger weight so final results reflect the true 35% share. If the survey also underrepresents a language community, you can include preferred language as a stratification variable and weight responses to better reflect the community’s presence.

Practical application

Post-stratification matters because many real-world surveys are non-random or suffer differential response rates across groups. Applying post-stratification makes estimates (means, totals, proportions) more representative of the population you care about, improving the credibility and usefulness of findings. For Hearo users, it’s especially valuable when some language communities or demographic groups respond less often: weighting responses by language, age, or area helps ensure the voices of underrepresented groups are not drowned out. Important caveats: it only corrects for variables you observe and have accurate population benchmarks for; it can increase estimate variance if some weights are very large; and it can't fix biases caused by measurement errors (for example, poor translation) or unobserved differences.

FAQ

How is post-stratification different from stratified sampling?

Stratified sampling is a design choice you make before data collection to ensure chosen groups are sampled in the right proportions. Post-stratification is performed after collection to adjust an already-obtained sample to match known population proportions. Put another way: stratified sampling controls the sample; post-stratification adjusts the sample.

Do I need accurate population data to use post-stratification?

Yes. Post-stratification requires reliable benchmark proportions for each stratum (for example census data or administrative records). If those benchmarks are wrong, the weighted results will be biased. Choose variables with trustworthy population margins.

Can post-stratification fix problems caused by poor translations or misunderstood questions?

No. Post-stratification can correct for unequal representation across observed groups, but it cannot fix measurement error caused by bad wording or mistranslation. That is why Hearo pairs automated translation with review and participant feedback: improved translations reduce measurement error, while post-stratification fixes representation.

Are there risks or limits to using post-stratification?

Yes. Large or highly variable weights can inflate the variance of estimates and make results unstable; very small sample sizes within strata make weighting unreliable; and post-stratification only adjusts for observed variables, not unknown sources of bias. Common mitigations are combining small strata, trimming extreme weights, or using raking to balance multiple margins.