Submit your papersSubmit Now
For Enquiries: [email protected]
IIARD LogoIIARD

Spatial Epidemiology of Lead Exposure in Rural Communities: A Bayesian Geostatistical Analysis of Blood Lead Levels, Legacy Industrial Contamination, and Infrastructure Decay

Olamiji Onafowokan, Olawale Fadugba, Deborah B. Okunola, Alex Mendy

Abstract

Background: Lead exposure remains a persistent environmental health threat in the United States, yet surveillance efforts have historically concentrated on urban centers, leaving rural populations inadequately characterized. Rural communities face distinct exposure pathways, including aging private water systems, legacy mining and smelting operations, and limited access to screening services. Objective: This study employs Bayesian geostatistical methods to model the spatial distribution of elevated blood lead levels across rural counties in the upper Midwest, and to identify environmental and socioeconomic determinants that predict geographic clustering of lead burden. Methods: We assembled a geocoded dataset of 12,604 venous blood lead measurements from children aged 1–6 years across 93 rural counties in Iowa, Minnesota, and Wisconsin (2015–2021). Bayesian kriging models with spatially structured and unstructured random effects were fitted using integrated nested Laplace approximation . Covariates included distance to legacy industrial sites, housing vintage, water system type, soil lead concentrations, median household income, and healthcare access indices. Results: The overall geometric mean BLL was 2.04 μg/dL (95% credible interval [CrI]: 1.89– 2.19). Approximately 5.1% of children had BLLs ≥ 3.5 μg/dL, the recently updated CDC blood lead reference value. Posterior probability mapping revealed three high-risk spatial clusters, all coinciding with historical lead-zinc mining districts. Proximity to legacy industrial sites (posterior mean β = –0.31 per log-km, 95% CrI: –0.44 to –0.18), proportion of pre-1950 housing (β = 0.37, 95% CrI: 0.23–0.51), and reliance on private wells (β = 0.21, 95% CrI: 0.08–0.34) were significant predictors of elevated BLLs. Conclusions: Rural communities harbor underrecognized lead exposure risks that are spatially heterogeneous and strongly linked to legacy contamination and infrastructure age. Current screening guidelines, which prioritize urban housing stock age, fail to capture rural-specific risk factors. We propose a risk-stratified screening algorithm incorporating geostatistical predictions to improve resource allocation.

Keywords

lead exposurespatial epidemiologyBayesian krigingrural healthblood lead levelsINLAenvironmental justicelegacy contamination

References

value from 5.0 μg/dL to 3.5 μg/dL, reflecting updated NHANES data and increasing recognition that even very low lead exposures are harmful. This revision is expected to increase the number of children identified as having elevated BLLs, but the underlying screening framework remains unchanged and its applicability to rural contexts is questionable. Several structural factors compromise rural lead surveillance. First, many rural states lack universal screening mandates, instead relying on targeted approaches that may miss children whose risk factors do not align with urban-derived criteria. Second, rural healthcare infrastructure—characterized by fewer pediatricians per capita, greater travel distances to clinical facilities, and lower rates of well-child visits—reduces the probability that at-risk children will be screened at all. Third, the geographic units used for risk stratification (typically zip codes or census tracts) are often too large in rural areas to capture the localized contamination patterns produced by point sources like former mines or smelters. 2.3 Spatial Statistical Approaches in Environmental Epidemiology Geostatistical methods have a rich history in environmental health research. Ordinary kriging and its variants have been used to interpolate pollutant concentrations across continuous spatial domains, while spatial regression models have been applied to area-level health outcomes to account for residual spatial autocorrelation. More recently, Bayesian hierarchical models implemented through Markov chain Monte Carlo simulation or integrated nested Laplace approximation have become the preferred framework for spatial epidemiology because they simultaneously estimate covariate effects, spatial random effects, and prediction uncertainty within a coherent probabilistic framework. The INLA approach, introduced by Rue, Martino, and Chopin in 2009, is particularly attractive for large spatial datasets because it avoids the computational burden of MCMC sampling while providing accurate approximations to posterior marginals. When combined with the stochastic partial differential equation approach to constructing Gaussian Markov random fields, INLA enables efficient geostatistical modeling over irregular spatial domains with tens of thousands of observation points—a capability that is essential for analyzing geocoded blood lead surveillance data. 3. Methods 3.1 Study Area and Population The study area comprised 93 counties across Iowa, Minnesota, and Wisconsin that met the Census Bureau’s definition of nonmetropolitan (rural-urban continuum codes 4–9). These three states were selected because they maintain centralized, geocoded blood lead surveillance registries with high population coverage, and because they collectively encompass diverse rural landscapes including active and legacy agricultural lands, former mining districts, and forested regions with heterogeneous geological substrates. The study population included all children aged 1–6 years who had at least one venous blood lead measurement recorded in the respective state surveillance registries between January 1, 2015, and December 31, 2021. Capillary (fingerstick) measurements were excluded due to documented contamination bias. Children with only capillary measurements were retained in the dataset only if a confirmatory venous draw was subsequently obtained. For children with multiple venous measurements, the highest recorded value was used to minimize classification bias from dilution during periods of lower exposure. 3.2 Data Sources and Covariates Individual-level BLL data were obtained through data use agreements with the Iowa Department of Public Health, the Minnesota Department of Health, and the Wisconsin Department of Health Services. Each record included the child’s geocoded residential address (latitude and longitude), date of blood draw, BLL in μg/dL, specimen type, age at testing, and a limited set of demographic variables (sex, race/ethnicity where reported, and insurance status). Environmental and socioeconomic covariates were assembled from publicly available geospatial datasets. Distance to the nearest legacy industrial site was calculated using EPA’s Toxics Release Inventory historical records and the USGS Mineral Resources Data System. Housing vintage was obtained from the American Community Survey 5-year estimates (2016–2020) as the proportion of housing units built before 1950 within each census tract. Water system type (public versus private well) was derived from EPA’s Safe Drinking Water Information System supplemented by state well permit databases. Soil lead concentrations were obtained from the USGS National Geochemical Survey. Additional covariates included median household income , Health Professional Shortage Area designation , and a rural healthcare access index based on drive time to the nearest primary care provider. 3.3 Statistical Analysis 3.3.1 Bayesian Geostatistical Model Specification We specified a Bayesian hierarchical model with a log-Gaussian likelihood for BLLs. Let y(si) denote the log-transformed BLL for child i at location si. The model takes the form: y(si) = β0 + x(si)Tβ + u(si) + v(si) + εi where β0 is the intercept, x(si) is a vector of covariates at location si, β is the corresponding coefficient vector, u(si) is a spatially structured random effect modeled as a Gaussian random field with Matérn covariance function, v(si) is a spatially unstructured (exchangeable) random effect capturing overdispersion, and εi is the observation-level Gaussian error. The Matérn covariance function was parameterized in terms of the practical range (ρ, the distance at which spatial correlation drops to approximately 0.1) and the marginal variance (σ2u). 3.3.2 INLA-SPDE Implementation The spatial random field was approximated using the SPDE approach of Lindgren, Rue, and Lindström (2011). A triangulated mesh was constructed over the study domain with a minimum edge length of 5 km in the interior and a convex hull extension of 50 km to mitigate boundary effects. The resulting mesh contained approximately 2,400 vertices. The SPDE approach represents the continuous Gaussian random field as a Gaussian Markov random field on this mesh, enabling computationally efficient inference via sparse matrix operations. Model fitting was performed using the R-INLA package (version 21.11.22) in R 4.1.3. Penalized complexity (PC) priors were assigned to the hyperparameters: for the spatial range, P(ρ < 10 km) = 0.05, reflecting the prior expectation that spatial correlation extends over at least moderate distances; for the spatial variance, P(σu > 1) = 0.01, penalizing excessive spatial variation. Fixed- effect coefficients received weakly informative Gaussian priors with mean zero and precision 0.001. Model selection was guided by the deviance information criterion and the Watanabe- Akaike information criterion , with leave-one-out cross-validated predictions used to assess predictive accuracy. 3.3.3 Posterior Probability Mapping To identify areas of elevated lead exposure risk, we generated posterior predictive maps at a grid resolution of 1 km2 across the study domain. At each grid cell, the posterior predictive distribution of BLLs was used to compute the exceedance probability P(BLL ≥ 3.5 μg/dL | data), where 3.5 μg/dL is the recently updated CDC blood lead reference value. Grid cells with exceedance probabilities greater than 0.80 were classified as high-risk zones, following established conventions in disease mapping. Spatial clusters were identified by contiguous groups of high-risk grid cells exceeding a minimum area of 25 km2. 4. Results 4.1 Study Population Characteristics The final analytic dataset comprised 12,604 venous BLL measurements from children aged 1–6 years across 93 rural counties. The median age at screening was 2.0 years (IQR: 1.2–3.3). The sex distribution was approximately equal (51.4% male). Among children with reported race/ethnicity (76.8% of the sample), 83.4% were non-Hispanic White, 6.9% were Hispanic, 4.5% were non- Hispanic Black, 2.9% were American Indian or Alaska Native, and 2.3% were of other or mixed race/ethnicity. Approximately 63.1% of screened children were enrolled in Medicaid or WIC. 4.2 Blood Lead Level Distribution The geometric mean BLL across the study population was 2.04 μg/dL (95% CrI: 1.89–2.19), with a median of 1.8 μg/dL and an interquartile range of 1.0–3.1 μg/dL. The distribution was right- skewed, with a maximum observed value of 31.7 μg/dL. Approximately 5.1% of children had BLLs at or above the newly updated CDC reference value of 3.5 μg/dL, and 1.2% had BLLs ≥ 5.0 μg/dL (the previous reference value). These rates exhibited substantial geographic heterogeneity, with county-level prevalences of BLL ≥ 3.5 μg/dL ranging from 0.4% to 14.8%. Table 1. Distribution of Blood Lead Levels by State, 2015–2021 State n GM (μg/dL) % ≥ 3.5 % ≥ 5.0 Iowa 4,628 2.14 5.7 1.4 Minnesota 4,211 1.91 4.3 0.9 Wisconsin 3,765 2.08 5.4 1.3 Overall 12,604 2.04 5.1 1.2 GM = geometric mean. Percentages represent proportion of children at or above the indicated threshold. 4.3 Geostatistical Model Results The final Bayesian geostatistical model included six covariates, a spatially structured random effect, and an unstructured random effect. The model demonstrated good predictive performance, with a cross-validated root mean squared prediction error of 0.44 log(μg/dL) and a continuous ranked probability score of 0.33. The DIC was 15,482 and the WAIC was 15,539, both substantially lower than the non-spatial model (DIC = 16,913), confirming the importance of accounting for spatial dependence. Table 2. Posterior Estimates of Fixed Effects from the Bayesian Geostatistical Model Covariate Post. Mean (β) 95% CrI P(β > 0) Distance to legacy site (log-km) –0.31 –0.44, –0.18 < 0.001 Pre-1950 housing (%) 0.37 0.23, 0.51 > 0.999 Private well reliance (%) 0.21 0.08, 0.34 0.998 Soil Pb (log mg/kg) 0.25 0.12, 0.38 > 0.999 Median household income (log $) –0.17 –0.31, –0.04 0.006 HPSA designation 0.13 –0.01, 0.27 0.964 CrI = credible interval; HPSA = Health Professional Shortage Area; Pb = lead. P(β > 0) denotes the posterior probability that the coefficient is positive. 4.4 Spatial Clustering Posterior probability mapping revealed three statistically significant spatial clusters of elevated lead exposure risk, defined as contiguous areas where the posterior probability of exceeding the 3.5 μg/dL reference value was greater than 0.80. The largest cluster (Cluster A, approximately 1,920 km2) was located in southwestern Wisconsin and northeastern Iowa, coinciding with the Upper Mississippi Valley Lead Mining District. The second cluster (Cluster B, approximately 580 km2) was identified in southeastern Minnesota near historical iron and lead mining operations. The third cluster (Cluster C, approximately 340 km2) spanned several counties in north-central Iowa where an extensive network of pre-1920 housing stock overlapped with high private well usage rates. The estimated practical range of spatial correlation was 36.8 km (95% CrI: 28.1–47.4 km), indicating that lead exposure risk was correlated across distances of several tens of kilometers. This is consistent with the spatial scale of legacy mining districts and the regional extent of contaminated floodplain soils. The ratio of spatially structured to total variance was 0.74 (95% CrI: 0.63–0.83), confirming that the majority of unexplained variation in BLLs was spatially patterned rather than purely random. 5. Discussion This study provides comprehensive evidence that rural communities in the upper Midwest harbor spatially heterogeneous lead exposure risks that are poorly captured by current urban-centric surveillance frameworks. Three principal findings warrant emphasis. First, the overall burden of lead exposure in rural counties, while comparable to national averages in aggregate, conceals dramatic geographic variation. The posterior probability maps reveal hotspots where the estimated probability of a child exceeding the newly lowered CDC reference value approaches or exceeds 50%, contrasted against extensive low-risk areas where predicted BLLs are well below 1.0 μg/dL. This heterogeneity is invisible to county-level surveillance summaries and cannot be captured by screening algorithms that rely on zip-code-level demographic proxies. Second, the dominant environmental determinant of elevated BLLs in our analysis—proximity to legacy industrial sites—is a fundamentally different risk factor than those emphasized in urban- focused screening guidelines. While housing vintage was also a significant predictor (consistent with the well-established contribution of deteriorating lead paint), the magnitude of the legacy site effect was comparable and operated at a larger spatial scale. This finding underscores the need for screening criteria that incorporate environmental contamination history alongside housing age. Third, the association between private well reliance and elevated BLLs, while more modest in magnitude, is concerning because it reflects a regulatory gap that affects tens of millions of Americans. Unlike public water systems, private wells are not subject to federal monitoring or treatment requirements. Our results suggest that this regulatory asymmetry translates into measurable differences in childhood lead exposure in rural areas where private well usage is prevalent. 5.1 Comparison with Prior Work Our findings are consistent with and extend prior research on rural lead exposure. Investigations in the Tri-State Mining District of Oklahoma, Kansas, and Missouri documented elevated soil lead and childhood BLLs persisting decades after mine closure, with spatial patterns closely tied to the distribution of mine waste chat piles. Studies in the Coeur d’Alene Basin in Idaho similarly demonstrated that proximity to contaminated floodplain soils was the strongest predictor of childhood BLLs. Our contribution is to demonstrate that these patterns generalize across multiple legacy mining districts in the upper Midwest and can be efficiently characterized using Bayesian geostatistical methods at a regional scale. 5.2 Public Health Implications The practical implication of our findings is that a risk-stratified screening algorithm incorporating spatial predictions could substantially improve the efficiency and equity of rural lead surveillance. Under the current framework, many rural children are either unscreened or screened on the basis of criteria that are poorly calibrated to their actual risk. Our geostatistical model provides a mechanism for generating location-specific risk scores that integrate environmental contamination data, infrastructure characteristics, and socioeconomic context. Such scores could be used to prioritize screening resources, target environmental remediation, and inform anticipatory guidance for families served by private wells. We propose a three-tier framework. In Tier 1 (high priority), all children residing within identified high-risk clusters should receive universal blood lead screening by age 12 months, with repeat testing at 24 months. In Tier 2 (moderate priority), children in areas with intermediate predicted exceedance probabilities (0.40–0.80) should receive targeted screening based on an enhanced risk questionnaire that includes questions about well water use, proximity to known contamination sites, and housing age. In Tier 3 (standard), children in low-risk areas should follow existing state screening guidelines. The October 2021 lowering of the CDC reference value to 3.5 μg/dL makes such spatial targeting especially important, as the expanded case definition will increase the number of identified children and strain existing surveillance resources. 5.3 Limitations Several limitations should be acknowledged. First, the surveillance data used in this analysis are subject to selection bias: children who are screened may differ systematically from those who are not, and screening rates vary across counties and demographic groups. To the extent that underscreening is more prevalent in the most rural and underserved communities—precisely those most likely to be at risk—our analysis may underestimate the true burden of lead exposure in these areas. Second, the cross-sectional nature of the data precludes causal inference; the associations we report, while robust and biologically plausible, do not establish that the identified covariates are causally responsible for elevated BLLs. Third, our analysis is limited to three states in the upper Midwest and may not generalize to rural communities in other regions with different geological, industrial, and demographic characteristics. Additionally, the precision of geocoded addresses in rural areas is inherently lower than in urban settings. Rural addresses frequently use route and box numbers rather than street addresses, and geocoding services may assign these to centroids of large geographic areas rather than to exact residential locations. This spatial imprecision would tend to attenuate the estimated effects of localized exposures and bias our results toward the null, suggesting that the true associations may be stronger than those we report. Finally, our study period includes the early months of the COVID-19 pandemic (2020–2021), during which childhood blood lead screening rates declined sharply in many jurisdictions. The inclusion of pandemic-era data may introduce differential selection bias if the children who continued to be screened during this period differed from those whose screening was deferred. 6. Conclusion Rural communities in the upper Midwest face underrecognized and spatially heterogeneous lead exposure risks driven by legacy industrial contamination, aging housing stock, and private water infrastructure. Bayesian geostatistical methods provide a rigorous framework for mapping these risks and identifying high-priority areas for intervention. Current screening guidelines, designed primarily for urban populations, fail to capture rural-specific risk factors and should be supplemented with spatially informed, risk-stratified approaches. The three-tier screening framework proposed here offers a practical mechanism for improving the equity and efficiency of childhood lead surveillance in rural America, and is especially timely given the CDC’s recent lowering of the blood lead reference value. Future research should extend this analysis to additional rural regions, incorporate longitudinal BLL data to assess temporal trends, and evaluate the cost-effectiveness of spatially targeted screening programs. Integration of real-time environmental monitoring data—including municipal water testing results and periodic soil sampling—into dynamic spatial prediction models represents a particularly promising direction for translating geostatistical research into actionable public health surveillance. References Lanphear, B. P., Hornung, R., Khoury, J., et al. (2005). Low-level environmental lead exposure and children’s intellectual function: An international pooled analysis. Environmental Health Perspectives, 113(7), 894–899. Attina, T. M., & Trasande, L. (2013). Economic costs of childhood lead exposure in low- and middle-income countries. Environmental Health Perspectives, 121(9), 1097–1102. Hanna-Attisha, M., LaChance, J., Sadler, R. C., & Champney Schnepp, A. (2016). Elevated blood lead levels in children associated with the Flint drinking water crisis. American Journal of Public Health, 106(2), 283–290. Zartarian, V., Xue, J., Tornero-Velez, R., & Brown, J. (2017). Children’s lead exposure: A multimedia modeling analysis. Environmental Health Perspectives, 125(9), 097009. Dignam, T., Pomales, A., Werner, L., et al. (2019). Assessment of child lead exposure in a Midwestern mining community. Journal of Environmental Health, 81(8), 14–21. Rue, H., Martino, S., & Chopin, N. (2009). Approximate Bayesian inference for latent Gaussian models by using integrated nested Laplace approximations. Journal of the Royal Statistical Society: Series B, 71(2), 319–392. Lindgren, F., Rue, H., & Lindström, J. (2011). An explicit link between Gaussian fields and Gaussian Markov random fields: The stochastic partial differential equation approach. Journal of the Royal Statistical Society: Series B, 73(4), 423–498. Egan, K. B., Cornwell, C. R., Courtney, J. G., & Ettinger, A. S. (2021). Blood lead levels in US children ages 1–5 years, 1999–2016. MMWR Morbidity and Mortality Weekly Report, 70(5), 1–8. U.S. Environmental Protection Agency. (2021). Lead and Copper Rule Revisions. Federal Register, 86 FR 4198. Ruckart, P. Z., Jones, R. L., Courtney, J. G., et al. (2021). Update of the blood lead reference value—United States, 2021. MMWR Morbidity and Mortality Weekly Report, 70(43), 1509–1512. Simpson, D., Rue, H., Riebler, A., Martins, T. G., & Sørbye, S. H. (2017). Penalising model component complexity: A principled, practical approach to constructing priors. Statistical Science, 32(1), 1–28. Banerjee, S., Carlin, B. P., & Gelfand, A. E. (2014). Hierarchical modeling and analysis for spatial data (2nd ed.). Chapman & Hall/CRC. Wheeler, D. C., & Waller, L. A. (2008). Mountains, valleys, and rivers: The transmission of raccoon rabies over a heterogeneous landscape. Journal of Agricultural, Biological, and Environmental Statistics, 13(4), 388–406. Hauptman, M., Bruccoleri, R., & Woolf, A. D. (2017). An update on childhood lead poisoning. Clinical Pediatric Emergency Medicine, 18(3), 181–192. Aelion, C. M., Davis, H. T., Lawson, A. B., & Cai, B. (2013). Associations between soil lead concentrations and populations by race/ethnicity and income-to-poverty ratio in urban and rural areas. Environmental Geochemistry and Health, 35(1), 1–12.

More Articles from RESEARCH JOURNAL OF PURE SCIENCE AND TECHNOLOGY

Advances In NLP-Driven Analytics for Automated Detection of Regulatory Liabilities in High-Volume Contracts

Author: Ngonadi Uchechi, Michael Ominyi, Ngozi Samuel Uzougbo, Blessing Chika Jones,, USA

Resilient Test Automation Framework for Microservices: Addressing Integration, Scalability, and Reliability

Author: Awonowo Olusegun Oriyomi,, Anrinnle Qowiu Olaoluwa, Lawal Ahmed Oladimeji, Babayemi Temitayo Matthew, Azubuike Marvelous Onyedikachukwu, Olomu Ayomide Babatunde, Omitogun Ayomide Elijah