A Novel Method for Generating Spatially Resolved Synthetic Populations for Health Impact Assessments in Vulnerable Populations

Geohealth. 2026 Jan 4;10(1):e2025GH001596. doi: 10.1029/2025GH001596. eCollection 2026 Jan.

Abstract

The spatial resolution of environmental exposure and sociodemographic population data is often mismatched given limited publicly available population data that complies with privacy requirements for individuals. To address this limitation, we developed a novel matching algorithm to construct a synthetic population at the address-level. To demonstrate how our approach can improve environmental justice (EJ) analyses and health impact assessments (HIAs), we examined sociodemographic patterns of residential proximity to major roadways in Greater Boston (Massachusetts) and HIA results, comparing our method with a random address allocation method. The synthetic population was developed at a census tract-level using US Census microdata and combinatorial optimization methods and then downscaled to address-level parcels by matching building attributes to synthetic households. We designated households within 50 m of a major road "high exposure" and households below state median household income "low income".We found misclassification for individual households (21% of the high exposure/low-income households in the matched data set were identified as such in the random allocation data set). We found modest aggregate differences in matched allocation (3.3% of low-income households had high exposure) compared to random allocation (3.4%). In a HIA, the difference between random and matched allocation would be stronger when there is a strong interactive effect between a sociodemographic effect modifier and exposure on the outcome. Address-level exposure assignment based on synthetic populations can provide more significant and nuanced health impact and EJ analyses. Our novel method can be applied to other regions of the US and expanded to other dimensions of population vulnerability.

Keywords: environmental justice; health impact assessment; synthetic population.

Plain language summary

Communities and decision makers often need to identify if there are disparities in the distribution of hazardous exposures and associated health outcomes. To do so requires understanding of both spatial patterns of exposures and of the attributes of exposed populations. While environmental exposure data are available at increasingly higher spatial resolution, data on high‐resolution population sociodemographic characteristics are limited by privacy requirements in the US. To support the investigation of environmental exposures and health outcomes across sociodemographic characteristics at address‐level resolution, we used publicly available US Census data to simulate an address‐level population with sociodemographic information. In a case study looking at proximity to major roadways in Greater Boston (Massachusetts), we compared exposure patterns between our approach and approaches where household attributes were not used for address assignment. We found large differences in how individual households were identified but modest differences in the percent of households identified as high‐exposure and low‐income. We also showed that differences in estimated health impacts would depend on whether there was a strong interaction between the environmental exposure and sociodemographic variable. The methods used to create the address‐level synthetic population can be replicated in other regions of the US using the same census data resources.