Quantifying Estimate Saturation through Mathematical Reliability Theory: An Application to Medical Scheme Member-Movement Data in South Africa, 2014–2024
Keywords:
Data Saturation, Reliability Theory, Spearman–Brown, Bootstrap, Relative Standard Error, Medical Schemes, Administrative Data, South AfricaAbstract
Background: “Data saturation” the point at which additional data cease to alter a study’s conclusions is widely invoked but rarely defined mathematically, leaving the question “how much data is enough?” to subjective judgement. We recast saturation as a property of estimator reliability and apply it to a large administrative dataset.
Methods: Reliability theory implies that the relative standard error (RSE) of an aggregate estimator decays as a square-root law, RSE(n) ≈ c/√n, the continuous analogue of the Spearman–Brown formula. Saturation is then defined as the sample size n* at which RSE first falls below a tolerance τ. Using Council for Medical Schemes member-movement records (2014–2024; 70 schemes; 26,959 usable stratum-year observations), we estimated a system turnover proportion and a dependant exit-share, characterised reliability decay by bootstrap resampling, and validated the square-root law.
Results: Pooled beneficiary turnover was 46.8% and the dependant share of exits 51.8%. Exit intensity rose monotonically with age from 16.4% in infants to 84.7% at 85+ years, a classical reliability wear-out profile. The bootstrap RSE followed the predicted law (c = 1.08), giving saturation at n* ≈ 467, 2,919 and 11,673 observations for tolerances of 5%, 2% and 1%. Saturation was parameter-specific: the dependant exit-share required ≈ 2,778 observations at 5%.
Conclusions: Framing saturation as a reliability threshold yields an explicit, reproducible stopping rule and a design tool for deciding how much administrative data is sufficient for a target precision.
