Survey Sample Size Calculator
Estimate a simple-random-sample size for a proportion using confidence level, margin of error, expected proportion, finite population, design effect and expected response rate. The page separates completed responses from invitations.
Set the survey precision assumptions
Treat this as proportion-survey planning under stated assumptions, then address frame quality, design, subgroup needs and non-response separately.
The proportion formula
For a large population, the initial sample size is z²p(1−p)/e². The z value represents the selected confidence level, p is the expected proportion, and e is the margin of error written as a decimal. The result estimates completed responses needed for a confidence interval around one proportion under simple random sampling.
At 95% confidence, p=0.5 and a five-percentage-point margin, the large-population result is about 385. The formula is not a rule that every survey with 385 responses is representative. Random selection, coverage, measurement and response behaviour remain essential.
Confidence level and margin of error
A higher confidence level requires a larger sample when other inputs remain fixed. A smaller margin of error also increases sample size rapidly because e is squared in the denominator. Halving the margin from five to 2.5 percentage points requires about four times the large-population sample, not twice as many.
The confidence level describes long-run performance of intervals under the sampling model. It does not mean there is a 95% probability that a fixed realised interval contains a fixed parameter after the data are observed. In ordinary reporting, explain the method without turning confidence into certainty.
Choosing an expected proportion
The product p(1−p) is largest at p=0.5, so 50% gives the most conservative sample size when the likely proportion is unknown. A reliable previous estimate can reduce the calculated number, but using an optimistic extreme merely to lower cost weakens planning. The estimate should match the same population, question and period as closely as possible.
Different survey questions can have different expected proportions. Plan for the key outcome that needs the greatest sample, or evaluate several values. Means, rates and rare-event detection use different sample-size formulas and are outside this proportion calculator.
Finite-population correction
When sampling without replacement from a known finite population N, the correction divides the initial size by 1 + (n₀−1)/N. It matters when the sample is a noticeable share of the population. For a very large or unknown population, enter zero and the initial result remains unchanged.
Population size means the number of eligible units in the defined frame, not the total population of South Africa unless every resident is eligible and reachable. Duplicate, outdated or missing frame records affect coverage. The correction assumes the frame and selection process are valid.
Design effect
Cluster sampling, unequal weights and other complex designs can produce more variance than a simple random sample of the same size. A design effect multiplies the corrected sample requirement to allow for that loss of efficiency. A value of 1 means no inflation. An appropriate value should come from prior comparable surveys or a statistician, not from a convenient guess.
Stratification can sometimes improve precision, while cluster similarity often increases required size. One overall design effect may not fit every estimate. South African household surveys with multistage sampling need design-specific analysis and appropriate variance estimation after collection.
Completed responses versus invitations
The completed-response result is divided by the expected response rate to estimate invitations. At 60% response, 370 completed surveys require planning for about 617 invitations. This assumes the invitations are eligible and response behaviour is close to the forecast. If the finite population has fewer members, the page caps invitations at the population size.
Sending more invitations can improve the count but does not automatically remove non-response bias. Respondents may differ systematically from non-respondents. Track disposition codes, reminders, contact modes and subgroup response rates, and avoid coercive or excessive follow-up.
Subgroups and reporting domains
An overall national sample can be large enough for the total but too small for provincial, age or language estimates. Each subgroup needs adequate completed responses for its own precision. Oversampling a small group can help, followed by weighting in the overall analysis, but weights affect variance and should be incorporated in design planning.
Do not divide the overall number evenly among many groups and assume the original margin still applies. Calculate priority domains explicitly and consider multiple-comparison and operational constraints. A statistician can balance precision, cost and design.
Survey error beyond sampling
Margin of error covers only sampling variability under the stated model. Leading questions, poor translation, interviewer effects, device access, frame gaps, duplicates, fraud and data processing errors can dominate the total error. A very large biased sample can give a precise estimate of the wrong population or construct.
Pilot the questionnaire, document selection, protect respondent privacy and predefine cleaning rules. Informed consent and lawful handling of personal information remain necessary regardless of sample size. Publish limitations with the estimate rather than using the calculated number as a quality badge.
Rounding the required count
Sample requirements are rounded up because part of a completed response cannot meet the target. The finite correction is applied before the design effect, and invitation planning follows the completed-response target. Rounding down at several intermediate stages can leave the realised design slightly below the stated assumptions.
If the target population itself is smaller than the inflated requirement, the output cannot exceed the whole population. Attempting a census does not guarantee every unit responds, so achievable precision and non-response plans may need reconsideration.
Monitoring fieldwork
Track valid completions against the sampling strata and not only the headline total. Early responses can overrepresent easy-to-contact groups. Releasing the remaining sample in controlled batches and monitoring disposition codes can improve balance, but any adaptive procedure should be documented.
Do not stop simply when the first required number arrives if selection probabilities or quotas have not been respected. Conversely, collecting far more convenience responses does not repair a missing probability design.
Precision for estimated differences
A sample planned for one proportion may not be large enough to detect a difference between two groups or two periods. Difference estimates use the uncertainty from both sides and require assumptions about the effect size, allocation and statistical test. Paired or repeated surveys add further structure.
If the primary question is whether provincial results differ, plan that comparison directly rather than relying on the national proportion calculation. Pre-specifying the main estimate prevents sample-size choices from being changed after results are visible.
Questions that affect this result
Why does 50% give the largest sample?
The variance term p(1−p) reaches its maximum at 0.5. Use 50% when no defensible prior proportion is available.
Does population size always change the sample a lot?
No. The finite-population correction becomes important when the planned sample is a meaningful fraction of the population. For very large populations the initial result changes little.
What does design effect mean?
It is a variance inflation relative to simple random sampling, often caused by clustering or unequal weights. Use evidence from a comparable design or statistical advice.
Are invitations the same as completed responses?
No. Invitations are inflated by the expected response rate. The precision formula applies to completed eligible responses under the sampling assumptions.
Does the margin of error include non-response bias?
No. It covers sampling variability only. Coverage, non-response, measurement and processing errors need separate prevention and assessment.