Reliable estimation of disease prevalence in small population subgroups is a common challenge in the social sciences. When disease-risk models are available, this issue necessitates reconstructing the joint dis-tribution of covariates, a challenging task when relying solely on survey data. The number of covariate combinations grows rapidly with the num-ber of covariates, making survey data scarce, and the estimated quanti-ties rarely match known marginal quantities available in aggregate form from other sources. We propose a flexible framework to reconstruct joint covariate distributions when survey data are sparse and reliable partial marginals are available from administrative sources. We match the joint covariate distribution when these are known. The missing components of the joint covariate distribution are then estimated sequentially, using highly expressive neural networks trained on survey data. Each covari-ate is predicted conditionally on those with known joint distributions, or on those previously reconstructed. Reliable marginals are incorporated directly into the loss function through strong penalization, thereby induc-ing reconstructed distributions to match external totals. We apply the framework to estimate the prevalence of diabetes and cardiovascular dis-ease in fine-grained U.S. subgroups. Combining Behavioral Risk Factor Surveillance System (BRFSS) data with Census marginals, we produce population and subgroup-specific prevalence estimates.

A Flexible AI Framework for Census Calibrated Small Domain Prevalence Estimation

Arletti, Alberto
;
Schiavon, Lorenzo
;
Stival, Mattia;Bertarelli, Gaia;Campostrini, Stefano
2026

Abstract

Reliable estimation of disease prevalence in small population subgroups is a common challenge in the social sciences. When disease-risk models are available, this issue necessitates reconstructing the joint dis-tribution of covariates, a challenging task when relying solely on survey data. The number of covariate combinations grows rapidly with the num-ber of covariates, making survey data scarce, and the estimated quanti-ties rarely match known marginal quantities available in aggregate form from other sources. We propose a flexible framework to reconstruct joint covariate distributions when survey data are sparse and reliable partial marginals are available from administrative sources. We match the joint covariate distribution when these are known. The missing components of the joint covariate distribution are then estimated sequentially, using highly expressive neural networks trained on survey data. Each covari-ate is predicted conditionally on those with known joint distributions, or on those previously reconstructed. Reliable marginals are incorporated directly into the loss function through strong penalization, thereby induc-ing reconstructed distributions to match external totals. We apply the framework to estimate the prevalence of diabetes and cardiovascular dis-ease in fine-grained U.S. subgroups. Combining Behavioral Risk Factor Surveillance System (BRFSS) data with Census marginals, we produce population and subgroup-specific prevalence estimates.
2026
Statistical Science: From Theory to Applied Research II. SIS-FENStatS 2026 2026
File in questo prodotto:
File Dimensione Formato  
2026_SISFENStat2026_ArlettiSchiavonStivalBertarelliCampostrini.pdf

non disponibili

Tipologia: Documento in Post-print
Licenza: Copyright dell'editore
Dimensione 560.64 kB
Formato Adobe PDF
560.64 kB Adobe PDF   Visualizza/Apri

I documenti in ARCA sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10278/5123831
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact