Reliable estimation of disease prevalence in small population subgroups is a common challenge in the social sciences. When disease-risk models are available, this issue necessitates reconstructing the joint dis-tribution of covariates, a challenging task when relying solely on survey data. The number of covariate combinations grows rapidly with the num-ber of covariates, making survey data scarce, and the estimated quanti-ties rarely match known marginal quantities available in aggregate form from other sources. We propose a flexible framework to reconstruct joint covariate distributions when survey data are sparse and reliable partial marginals are available from administrative sources. We match the joint covariate distribution when these are known. The missing components of the joint covariate distribution are then estimated sequentially, using highly expressive neural networks trained on survey data. Each covari-ate is predicted conditionally on those with known joint distributions, or on those previously reconstructed. Reliable marginals are incorporated directly into the loss function through strong penalization, thereby induc-ing reconstructed distributions to match external totals. We apply the framework to estimate the prevalence of diabetes and cardiovascular dis-ease in fine-grained U.S. subgroups. Combining Behavioral Risk Factor Surveillance System (BRFSS) data with Census marginals, we produce population and subgroup-specific prevalence estimates.
A Flexible AI Framework for Census Calibrated Small Domain Prevalence Estimation
Arletti, Alberto
;Schiavon, Lorenzo
;Stival, Mattia;Bertarelli, Gaia;Campostrini, Stefano
2026
Abstract
Reliable estimation of disease prevalence in small population subgroups is a common challenge in the social sciences. When disease-risk models are available, this issue necessitates reconstructing the joint dis-tribution of covariates, a challenging task when relying solely on survey data. The number of covariate combinations grows rapidly with the num-ber of covariates, making survey data scarce, and the estimated quanti-ties rarely match known marginal quantities available in aggregate form from other sources. We propose a flexible framework to reconstruct joint covariate distributions when survey data are sparse and reliable partial marginals are available from administrative sources. We match the joint covariate distribution when these are known. The missing components of the joint covariate distribution are then estimated sequentially, using highly expressive neural networks trained on survey data. Each covari-ate is predicted conditionally on those with known joint distributions, or on those previously reconstructed. Reliable marginals are incorporated directly into the loss function through strong penalization, thereby induc-ing reconstructed distributions to match external totals. We apply the framework to estimate the prevalence of diabetes and cardiovascular dis-ease in fine-grained U.S. subgroups. Combining Behavioral Risk Factor Surveillance System (BRFSS) data with Census marginals, we produce population and subgroup-specific prevalence estimates.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_SISFENStat2026_ArlettiSchiavonStivalBertarelliCampostrini.pdf
non disponibili
Tipologia:
Documento in Post-print
Licenza:
Copyright dell'editore
Dimensione
560.64 kB
Formato
Adobe PDF
|
560.64 kB | Adobe PDF | Visualizza/Apri |
I documenti in ARCA sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



