2. Materials and Methods
Kisumu is the third-largest city in Kenya located in the Lake Victoria area in the former Nyanza Province. It is the second-largest city in the Lake Victoria Basin after Kampala. The city has a population of 438,588 according to the recent 2026 population estimates. The metro region, including Maseno and Ahero, had a population of 6,771 (Maseno) and 11,801 (Ahero) people respectively, totaling to a combined population of 457,160 people, according to the 2019 Kenya Population and Housing Census, which was conducted by the Kenya National Bureau of Statistics. Kisumu's elevation is 1,131 m (3,711 ft) and is about 320 km (200 mi) northwest of Nairobi, on the shores of Lake Victoria. It lies at the northeastern edge of the Winam Gulf, a long, shallow arm that protrudes from the main body of Lake Victoria
| [9] | Machani, M. G., Onyango, S. A., Nzioki, I., Milanoi, S., Nattoh, G., Githure, J., Atieli, H., Wang, C., Lee, M., Zhou, G., Githeko, A., Afrane, Y. A., Ochomo, E., & Yan, G. (2024). Bionomics and distribution of malaria vectors in Kisumu city, Western Kenya: Implications for urban malaria transmission. Research Square. https://doi.org/10.21203/rs.3.rs-4943539/v1 |
[9]
. Kisumu is 24 km (15 mi) south of the equator and has, due to its elevation, moderate temperatures. Kisumu features a tropical rainforest climate with no true dry season and significant rainfall year-round. January is the driest month, while the month of April receives the most rainfall. Kisumu has highly fertile land and variations in temperature and rainfall, with two rainy seasons per year across the region providing a suitable environment for a broad range of agricultural crops. The Kisumu region has approximately 1.6 million hectares of agricultural land. However, it is estimated that only 58 percent of the land is currently used. The majority of farming in the lake basin region is subsistence agriculture, leading to relatively low production volume
| [9] | Machani, M. G., Onyango, S. A., Nzioki, I., Milanoi, S., Nattoh, G., Githure, J., Atieli, H., Wang, C., Lee, M., Zhou, G., Githeko, A., Afrane, Y. A., Ochomo, E., & Yan, G. (2024). Bionomics and distribution of malaria vectors in Kisumu city, Western Kenya: Implications for urban malaria transmission. Research Square. https://doi.org/10.21203/rs.3.rs-4943539/v1 |
[9]
.
2.1. Study Site Map of Kisumu
2.2. Satellite Data
We created a LULC map of Kisumu using the 10-meter visible and near-IR (infrared) imagery. We first acquired cloud-free Level-2A surface reflectance data from the European Space Agency covering the Kisumu metropolitan region and surrounding agricultural landscape near Lake Victoria. The key 10-m bands—blue (B2), green (B3), red (B4), and near-infrared (IR) (B8)—are stacked and preprocessed with atmospheric correction, cloud masking, and clipping to the study boundary. (
Table 1)
Table 1. Sentinel-2 with spectral bands at different spatial resolutions.
Band Number | Sentinel-2A | Sentinel-2B | Spatial resolution (m) |
Central wavelength (nm) | Bandwidth (nm) | Central wavelength (nm) | Bandwidth (nm) |
1 | 442.7 | 20 | 442.3 | 20 | 60 |
2 | 492.7 | 65 | 492.3 | 65 | 10 |
3 | 559.8 | 35 | 558.9 | 35 | 10 |
4 | 664.6 | 30 | 664.9 | 31 | 10 |
5 | 704.1 | 14 | 703.8 | 15 | 20 |
6 | 740.5 | 14 | 739.1 | 13 | 20 |
7 | 782.8 | 19 | 779.7 | 19 | 20 |
8 | 832.8 | 105 | 832.9 | 104 | 10 |
8a | 864.7 | 21 | 864.0 | 21 | 20 |
9 | 945.1 | 19 | 943.2 | 20 | 60 |
10 | 1373.5 | 29 | 1376.9 | 29 | 60 |
11 | 1613.7 | 90 | 1610.4 | 94 | 20 |
12 | 2202.4 | 174 | 2185.7 | 184 | 20 |
Training samples representing urban surfaces (dense buildings and roads in central Kisumu), peri-urban transition zones with mixed housing and vegetation, and rural farmland dominated by cropland and smallholder fields were collected from high-resolution reference imagery using field entomological sampling data. A supervised classification algorithm (i.e., maximum likelihood) was applied to classify the stacked spectral bands and indices into the three land-cover categories. Accuracy assessment included remotely validating the predicted georeferenced capture points to evaluate the reliability of the map and post-processing (such as smoothing and removing isolated pixels). This rendered a final 10-m resolution land-use map distinguishing urban, peri-urban, and rural farmland zones. We assumed these
An. stephensi regression maps would enable implementing a S&D-LSM program. The LULC-classified map revealed the capture points and surface area (m
2) of each site, including remote, sampled, georeferenced, stratified sentinel sites, as well as potential
An. stephensi habitats (
Figure 2).
Figure 2. LULC Map of Kisumu County.
2.3. Regression Model
Initially, a Poisson regression with statistical significance was calculated at a 95% confidence level in R. The Poisson process in our analysis was provided by the limit of a binomial distribution of the larval sampled
An. stephensi habitat estimator determinants.
Table 1 reveals the habitat covariates used in the regression model.
(1)
We viewed the distribution as a function of the expected number of the estimator determinant, stratified capture point data in Kisumu County, using the sample size
n to quantify the fixed p in Equations (
2-6):
(2)
Based on the sample size n, the distribution was solved using as
(4)
(5)
We generated plots and calculated the Poisson distribution, which was also used to fit a generalized linear mixed model (GLMM) to the stratified
An. stephensi immature habitat data capture points, by maximum likelihood estimation (MLE) of the parameter vector β. While linear regression assumes a straight-line relationship and normally distributed outcomes, GLMMs can handle a variety of data types and distributions, making them suitable for real-world scenarios where data often deviates from these assumptions
. The procedure estimated the sampled
An. stephensi's larval habitat stratified capture point parameters numerically through an iterative fitting process. The dispersion parameter was approximated by the quantified residual deviance and by Pearson’s chi-square divided by the degrees of freedom (d. f.). Pearson's chi-squared test is a statistical test applied to sets of categorical data to evaluate how likely it is that any observed difference between the sets arose by chance, which is the most widely used of many chi-squared tests
| [13] | Greenwood, P. E. and Nikulin, M. S. A Guide to Chi-Squared Testing. 1st Edition. New York, NY: John Wiley & Sons, 1996. |
[13]
.
Covariances, standard errors, and p-values were subsequently computed using the determinants of the stratified capture point estimators of the sampled
An. stephensi habitats, based on the asymptotic normality derived from MLE. We first wrote the probability density function of the Poisson distribution,
. Subsequently we wrote the likelihood function, which was simply the product of the probability density function (PDF) for the observed sampled discrete, integer values x
1, …, x
n of the stratified
An. stephensi habitat covariates. We then calculated and generated the natural log likelihood function using Equations (
7-11).
(7)
(8)
(9)
(10)
(11)
We calculated the derivative of the natural log likelihood function with respect to λ in the Poisson model. Subsequently, we calculated the derivative of the natural log likelihood function with respect to the parameter λ in the h capture point abitatmodel using Equations (
12-15).
(12)
(13)
Thereafter, we set the derivative equal to zero and solved for λ in the model. Lastly, we set the derivative in the previous step equal to zero and simply solved for λ, which revealed:
Hence, the MLE became , which was equivalent to the sample mean of the dataset of the stratified capture point estimator determinants from the n sampled An. stephensi habitats. Next, we wrote the likelihood function. This simply was a product of the PDF for the observed values x1, …, xn.
It is important to note that the sample size
n completely dropped out of the probability function, which in this experiment, had the same functional form for all the stratified capture points and the discrete integer values of the sampled
An. stephensi habitat estimator determinants. As expected, the Poisson distribution was normalized so that the sum of probabilities equaled 1. The ratio of probabilities was determinable by
, which was subsequently expressible as Equation (
16).
(16)
2.4. Spectral Signature Interpolation
Spectral capture point features describing the land surface environment were extracted from multispectral satellite imagery obtained from Sentinel-2. After preprocessing steps including atmospheric correction, cloud masking, and raster alignment, spectral indices were calculated to capture environmental characteristics relevant to human settlement patterns and built infrastructure. These spectral indices, along with raw spectral band values, formed the environmental feature set associated with each georeferenced capture point An. stephensi habitat.
To interpolate the spectral signature relationships from a An. stephensi's habitat capture point to county-wide coverage, supervised machine learning algorithms were implemented in Python using the Scikit-learn and XGBoost libraries. The feature matrix, , consisted of spectral bands, derived spectral indices, and spatial coordinates, while the target variable, , corresponded to the predicted Poisson event intensity .
2.5. Geospatial and Machine Learning Libraries
2.5.1. Machine Learning-Based Spectral Signature Interpolation
Random Forest regression was implemented as an ensemble tree-learning algorithm that combined predictions from multiple decision trees to improve predictive performance. We averaged their predictions to obtain a robust capture point estimate. For an input vector x, the Random Forest prediction was defined employing Equations (
17 and
18).
where:
T = number of decision trees
= prediction from the -th tree
Each tree was trained on a bootstrap sample of the training dataset, and node splits were selected by minimizing the mean squared error (MSE):
2.5.2. Gradient Boosting Regression
Gradient Boosting Regression was also implemented to sequentially improve prediction accuracy. Using Gradient boosting, the predictive larval
An. stephensi's habitat capture point model was expressible using Equation (
19-20).
where:
= iteration
= a regression tree
= learning rate parameter
At each iteration, the algorithm fits a new tree to the residual errors of the previous capture point model, thus minimizing the loss function through gradient descent.
2.5.3. Extreme Gradient Boosting (XGBoost)
To further improve predictive performance, Extreme Gradient Boosting (XGBoost) was applied. XGBoost enhanced gradient boosting by incorporating regularization and efficient tree optimization. The objective function minimized by XGBoost was:
(20)
where:
= loss function (squared error)
= regression trees
= regularization term controlling model complexity
The regularization function was defined in the habitat model as:
where:
= number of leaves in the tree
= leaf weights
and = regularization parameters
2.6. Spatial Prediction for Kisumu
Once the machine learning models were trained using the
georeferenced larval
An. stephensi habitat capture points in Kisumu, the fitted models were applied to raster layers containing spectral features. For each raster cell,
, The trained model produced a predicted event intensity, where
represented the trained machine learning model using Equation (
21):
This process generated a continuous spatial surface representing predicted An. stephensi's larval habitat breeding sites across the county.
2.7. Spatial Cluster Detection
To identify spatial clusters of predicted activity, local spatial autocorrelation statistics were applied to the predicted raster surface. The Getis–Ord
statistic (Equation (
22)) was used to detect hotspots and cold spots:
(23)
where:
= predicted intensity value at location
= spatial weight between locations and
= global mean of predicted values
= standard deviation
High positive Values indicated statistically significant hotspots, while negative values indicated cold spots.
Next, a model was evaluated among the stratified capture point
An. stephensi habitat covariates, georeferenced in Kisumu at the county-level intervention sites, using Moran’s
I. In statistics, Moran's
I is a measure of spatial eigen-autocorrelation
| [14] | Griffith, 2003 Griffith, D. A. Spatial Autocorrelation and Spatial Filtering: Gaining Understanding Through Theory and Scientific Visualization. 1st Edition. New York, NY: Springer-Verlag Berlin Heidelberg; 2023.
https://doi.org/10.1007/978-3-540-24806-4 |
[14]
. We employed PySAL to compute Moran's
I to detect the
An. stephensi larval habitat estimator determinant, stratified, hot and cold, autocorrelated, capture point geolocations.
We employed an autoregressive conditional weight matrix that represented the degree of relationship between the georeferenced
An. stephensi habitat, capture points for which we verified latent autocorrelation coefficients. We examined the residual autocovariance, which was arranged like a mosaic, and computed the weight matrix using the Contiguity-Based Weights method in PySAL. Contiguity-based weights are usable in eigen-spatial analysis to define neighboring observations by creating weight matrices, which express the connectivity between observations
| [14] | Griffith, 2003 Griffith, D. A. Spatial Autocorrelation and Spatial Filtering: Gaining Understanding Through Theory and Scientific Visualization. 1st Edition. New York, NY: Springer-Verlag Berlin Heidelberg; 2023.
https://doi.org/10.1007/978-3-540-24806-4 |
[14]
. We computed Moran's
I for a two-dimensional dataset matrix, Z, using:
from libpysal.weights import lat2W
from esda.moran import Moran
import numpy as np
Z = np.random.rand(200,150)
# Create the matrix of weights’
w = lat2W(Z.shape[0], Z.shape[1])
# Create the pysal Moran object
mi = Moran(Z, w)
# Verify Moran's I results
print(mi.I)
print(mi.p_norm)
We started with a random matrix, Z. In our instance, the outcome of Moran's I approximated 0, serving as an effective test. Once we verified that it worked efficiently with the random Z data, we substituted the regressed model forecasts with the stratified, county-level, sampled, capture point data. An eigen-autocorrelation model specification was subsequently employed to describe the Poisson, random, stratified entomologic data. The resulting model specification took on the following form (1) where μ was the scalar conditional mean of Y, and ε was an n-by-1 error vector whose elements were statistically independent and identically distributed (i.i.d.) normally distributed random variates.
The spatial covariance matrix for equation (
1) employed the stratified, covariates in E [(Y - μ l)' (Y - μ l)] = Σ = [(I - ρ W')(I - ρ W)]-1σ2, where E (●) denoted the calculus of expectations, I was the n-by-n identity matrix delineating the matrix transpose operation, and σ
2 was the error variance. Our assumption was that a mixture of latent, positive and negative eigenized autocorrelated coefficients may be present in the stratified
, capture point,
An. stephensi's larval
habitat model estimator determinants, which in turn, we assumed would provide an even more explicit representation of the count variable Poisson regression model results. Varying autoregressive parameters appeared in the covariance matrix, which, for our regression model specification, was describable employing:
The diagonal matrix of the autoregressive parameters, <ρ >diag, contained two sampled parameters: ρ+, for those stratified
An. stephensi habitat covariates displaying positive eigen-spatial dependency, and ρ, for those displaying negative eigen-spatial dependency. Since ρ+ = 0 (and hence I+ = 0 and I- = I) or ρ- = 0 (and hence I- = 0 and I+ = I), equation (
24) reduces to equation (
1)
| [15] | Jacob, B. G., Izureta, R., Bell, J., Parikh, J., Louh, D., Gates, T., Murray, K., White, L., and Aceng, J. R. Approximating Non-Asymptoticalness, Skew Heterr (2023).
https://doi.org/10.5923/j.ajms.20231301.01 |
[15]
. This explanatory indicator variable classification was made in accordance with the quadrants of the corresponding Moran scatterplot generated using the, stratified, “Poissonized’ covariates. To identify county, georeferenceable, larval
An. stephensi's habitat stratified clusters, Thiessen polygon surface partitioning was subsequently generated using a Python 2.x script to construct geographic neighbor matrices. These were subsequently employed in the estimator determinant, eigen-autocorrelation, prognosticative, capture point stratified analysis. We compared the performance of non-hybrid neighborhood matrices, specifically contiguity-based, distance-sensitive, covariate-based (considering socioeconomic and sociodemographic characteristics), to predictively model the relative risk of small LULC-classified areas geolocated in Kisumu.
Neighbors based on contiguity were constructed by assuming that neighbors of a given georeferenced capture point, LULC-classified area in Kisumu at the county level, shared a common boundary. Neighbors can be of type Queen if a single shared boundary point meets the contiguity condition, or Rook if more than one shared point is required to meet the contiguity condition (14). The function poly2nb() of the spdep package was used to construct a list of stratifiable neighbors in Kisumu based on locations with contiguous boundaries, that is, georeferenced areas sharing one or more stratified capture points. We employed poly2nb() to calculate the neighbors of each of the regions of each intervention capture point site based on Queen contiguity. The default type in poly2nb () is queen = TRUE, so neighbors of a given area are other areas (14).
Next, a spatially autoregressive [SAR] model specification was subsequently employed to describe the autoregressive variance in the aggregation/non-aggregation-oriented, stratified, georeferenced, larval An. stephensi habitat, capture point, estimator determinants. An eigen-spatial filter model capture point signature specification was employed to describe both heterogeneous, Gaussian and Poisson, random, stratifiable potential, and estimator determinant effects. The resulting SAR model specification took on the following form,
In PySAL, where μ was the scalar conditional mean of Y, and ε was an n-by-1 error vector whose elements were normally
independent and identically distributed (i.d.d) covariates. The spatial covariance matrix for equation (
25), fit the stratified i.d.d. covariates using E [(Y - μl)' (Y - μl)] = Σ = [(I - ρ W') (I - ρ W)]
-1σ
2, where E (●) denoted the calculus of expectations, I was the n-by-n identity matrix denoting the matrix transpose operation, and σ
2 was the error variance. When a mixture of positive and negative latent eigen-autocorrelation is present in a non-time series, dependent, autoregressive vulnerability model, a more explicit representation of both effects leads to a more accurate interpretation of empirical results (14). Alternatively, the excluded values were set to zero, and as such, the mean and variance had to be adjusted in the model.
The model specification was subsequently transformed to
(26)
, where the diagonal matrix of the stratified An. stephensi larval habitat capture point sampled parameters, < ρ >diag, contained the uncertainty-oriented autoregressive variables ρ+ for those georeferenced estimator determinants displaying positive eigen-spatial dependency, and ρ for those displaying negative eigen-spatial dependency. For instance, by letting σ2 = 1 and employing a 2-by-2 regular square tessellation, the habitat model outcome enabled positing a positive relationship between the grid-stratified covariates when y1 and y2 had a negative relationship between the covariates, y3 and y4, and no relationship between covariates y1 and y3 and between y2 and y4. This covariance specification yielded:
(27)
where I
+ was a binary 0-1 explanatory measurement indicator, county, entomological stratified prognosticative variable. The specification also denoted those sampled, uncertainty-free, normalized, exogenous, regressed predictors displaying positive eigen-spatial dependency when I
- was a binary 0-1 variable, whilst denoting those estimator determinants displaying negative eigen-spatial dependency, employing I
+ + I
- = 1. If either ρ
+ = 0 (and hence I
+ = 0 and I
- = I) or ρ
- = 0 (and hence I
- = 0 and I
+ = I), then equation (
27) reduced to equation (
26). This variable indicator classification was made in accordance with the quadrants of the corresponding Moran scatterplot created using the stratified, aggregation/non-aggregation-oriented, sampled,
An. stephensi habitat capture point diagnostic estimator determinants in PySAL.
If positive and negative eigen-autocorrelation processes counterbalance each other in a mixture, the sum of the two eigen-autocorrelation parameters--(ρ
+ + ρ.) will be close to 0
| [14] | Griffith, 2003 Griffith, D. A. Spatial Autocorrelation and Spatial Filtering: Gaining Understanding Through Theory and Scientific Visualization. 1st Edition. New York, NY: Springer-Verlag Berlin Heidelberg; 2023.
https://doi.org/10.1007/978-3-540-24806-4 |
[14]
. Here, Jacobian estimation was implementable by utilizing the non-homogeneous, diagnostic, indicator values derived from the eigen-spatial filter georeferenced, aggregation/non-aggregation-oriented, stratified, exogenous
An. stephensi larval sampled habitat variables (I
+ - γ I
-) in eigenvector eigen-geospace, which required estimating ρ
+ and γ with ML techniques, and setting
. An empirical eigen-autocorrelation weight matrix was generated for quantitating the autocovariance of the non-time series, dependent, potentially non-homoscedastic, multicollinear, zero autocorrelated, non-Gaussian, aggregation/non-aggregation-oriented, sampled, asymptotical, stratified, georeferenced, capture point, habitat estimator determinants, which consisted of finding the normalized vectors u
i stored as columns in the matrix U = [u
1 ⋯ u
n]. This satisfied Λ = diag (λ
1 ⋯ λ
n),
and
for i ≠ j. It is important to note that double centering of Ω implied that the eigen-orthogonal eigenvectors rendered from the eigen-decomposed, georeferenced, stratified habitat regressors were centered, and at least one eigenvalue was equal to zero. Introducing these eigenvectors in the original formulation of Moran's
I led to:
(28)
The autocovariance is a function that gives the covariance of the process with itself at pairs of capture points, which is closely related to the eigen-autocorrelation (14). We centered vector z = Hx and employed the properties of idempotence of H, which subsequently made the pcounty-level, hot and cold spot, eigen-autocorrelation forecast models equivalent to
(29)
As the eigenvectors u
i and the vector z were centered in the georeferenced, aggregation/non-aggregation-oriented, vulnerability-oriented,
An. stephensi habitat, capture point hot and cold spot, regression model forecasts, equation (
29) was rewritten:
(30)
Equation (
30) where n was the number of null eigenvalues of Ω (r ≥ 1). These eigenvalues and corresponding eigenvectors were removed from Λ and U, respectively. Equation (
30) was strictly equivalent to
(31)
Moreover, it was demonstrated that Moran's I for a given georeferenced capture point, An. stephensi habitat eigen-spatial filter eigenvector ui was equal to I(ui) = (n/1T W1) λi, so the equation was rewritten as:
(32)
The term cor2 (ui, z) represented the part of the variance of z that was explainable by ui in the, hot/cold spot, prognosticative, capture point, entomological regression model when z = β i ui+ ei. Since the eigenvectors ui were eigen-orthogonalized, capture point habitat regression coefficients of the linear models z = β i ui+ ei were derivable from the r model z = Uβ + ε = β iui + ⋯ + β n-r un-r + ε.
The maximum value of 1 was quantifiable by all the variations of z, as parsimoniously expounded by the eigenvector u
1, which corresponded to the highest eigenvalue λ
1 in the weighted autocorrelation, non-time series dependent uncertainty-oriented matrix constructed from the stratified estimator determinants. Here, cor
2 (u
i, z) = 1 (and cor
2 (u
i, z) = 0 for i ≠ 1), and the maximum value of I was intuitively deducible for Equation (
31), which was equal to I
max = λ
1(n/1
TW1). The minimum value of I in the error matrix was obtainable as with all the variations of z, which in this experiment was definable by the eigenvector u
n-r corresponding to the lowest eigenvalue λ
n-r model forecast. This minimum value was equal to I
min = λ
n-r (n/1
TW1). If the sampled renderings were not definable due to the presence of quadratic heteroscedasticity, latent multicollinearity, and/or asymptoticalness in the model forecasts, the part of the variance explained by each eigenvector was equal, on average, to cor
2 (u
i, z) = 1/n-1. Because the diagnostic, epidemiological, sampled, county-level, and prognosticated capture point,
An. stephensi's habitat variables in z were randomly permuted; it was assumed that we would obtain this result.
The observed and expected numbers of stratified, sub-county-level, An. stephensi habitat capture point, estimator determinants were denoted by Y = (Y1, …, Yn) and E = (E1, …, En), respectively. In addition, we denoted the matrix of p [i.e., sampled breeding site covariates]. β = (β0, β1, …, βp) denoted the vector of the covariate effects, while Rk represented the risk at a georeferenced capture point geolocation.
2.8. Vector Host Population Buffer Zones
Figure 3. An. stephensi flight distance measurements.
Figure 4. Potential An. stephensi habitat.
3. Results
A Poisson probability regression model was determined by a 95% confidence level, which was employed to ascertain whether the proportions of stratified, sampled An. stephensi's habitat, capture point, estimator determinants differed by sub-county-level model predictions. We considered inference, hypothesis testing, and modeling causal relationships amongst the larval An. stephensi's habitat sampled estimator determinants. The regression analyses assumed independent counts (i.e., ni), taken at the georeferenced county locations i = 1, 2... n, where each of the regressed discrete integer values was from a Poisson distribution. These counts were described by a set of explanatory variables denoted by matrix Xi, a 1×p vector of estimator determinants for a sampled capture point sub-county geolocation i. The expected value of these data was given by where the vector of non-redundant parameters was the Poisson rates parameter which was given by . The rates parameter λi (Xi) was both the mean and the variance of the Poisson distribution, whereby each stratified, sampled, sub-county-level entomological estimator determinant was denoted as i. The dependent variable was the prevalence of malaria mortality rates in Kisumu in the regression model. The analysis was performed in R. The sampled data were log-transformed before analyses to normalize the distribution and minimize standard error. All of the covariate estimates for the model were tested using a backward stepwise regression. The most significant predictor was children.
Table 2. Output Summary Poisson Table.
Coefficients | Estimate | Standard Deviation Error | Z Value | Pr (>|z|) |
Intercept | -2.121858 | 1.509615 | -1.412 | 0.158 |
Children <5 | -0.001897 | 0.001801 | -1.054 | 0.292 |
A spectral signature capture point was performed in PySAL using the georeferenced, stratified regressed capture point, larval, An. stephensi habitat stratified, estimator determinants. An analogous variant of conventional correlation is serial correlation, which pertains to the correlation between values for observations of a single variable, according to some delivery of those values (14). Each georeferenced capture point location became a feature vector containing predictors such as population density and demographic characteristics. The response variable represented a composite larval, An. stephensi habitat exposure index. Geographic coordinates of the centroid were included so spatial patterns could be learned implicitly by the models or explicitly augmented with spatial lag features (e.g., neighboring capture point sentinel site averages).
Using Random Forest, Gradient Boosting, and XGBoost models, we trained supervised learning algorithms to map the relationship between predictor variables and malaria-related outcomes across Kisumu County. This created a structured dataset suitable for comparative machine-learning-based spatial risk prediction.
Table 3. Summary Statistics of Covariates.
Variable | Mean | Std. Dev. | Min | Max |
Distance to Water (m) | 248.5 | 198.4 | 10 | 850 |
Elevation (m) | 1152.3 | 18.7 | 1120 | 1185 |
Land Surface Temp (°C) | 29.8 | 2.1 | 25.8 | 34.5 |
Rainfall (mm) | 118.4 | 34.6 | 60 | 180 |
Population Density (persons/km²) | 1035 | 412 | 280 | 2150 |
Built-up Index | 0.56 | 0.18 | 0.22 | 0.89 |
Table 4. Variable Importance from Random Forest Model.
Rank | Predictor Variable | Importance Score |
1 | Distance to Water | 0.241 |
2 | Land Surface Temperature | 0.212 |
3 | Built-up Index | 0.176 |
4 | Rainfall | 0.141 |
5 | Population Density | 0.074 |
6 | Elevation | 0.041 |
Table 5. Variable Importance from XGBoost.
Rank | Predictor Variable | Gain (%) |
1 | Distance to Water | 24.8 |
2 | Built-up Index | 20.5 |
3 | Land Surface Temperature | 18.7 |
4 | Rainfall | 14.9 |
5 | Population Density | 7.1 |
6 | Elevation | 3.2 |
Table 6. Model Performance Comparison.
Model | Accuracy | Precision | Recall | F1-Score | AUC-ROC |
Random Forest | 0.86 | 0.84 | 0.82 | 0.83 | 0.91 |
Gradient Boosting | 0.88 | 0.86 | 0.85 | 0.85 | 0.93 |
XGBoost | 0.91 | 0.89 | 0.88 | 0.89 | 0.96 |
Table 7. Habitat Suitability Prediction by XGBoost.
Habitat ID | Predicted Probability | Predicted Class |
H001 | 0.92 | Suitable |
H002 | 0.18 | Unsuitable |
H003 | 0.87 | Suitable |
H004 | 0.09 | Unsuitable |
H005 | 0.95 | Suitable |
H006 | 0.31 | Unsuitable |
The machine learning models identified proximity to water bodies, land surface temperature, urbanization level (built-up index), and rainfall as the strongest determinants of Anopheles stephensi larval habitat occurrence. Among the evaluated algorithms, XGBoost achieved the highest predictive performance (AUC = 0.96), demonstrating superior ability to quantify complex nonlinear relationships and interactions among environmental and anthropogenic predictors. These findings suggest that urban water sources and warm, densely populated environments may provide favorable conditions for An. stephensi establishment across Kisumu County,
During training, the entomological dataset was divided into training and validation sets, allowing the model to learn complex anthropogenic interactions among estimator determinants such as socioeconomic and sociodemographic conditions. Once predicted risk values were quantified spatial statistical techniques were applied to detect hotspot and cold-spot clusters of An. stephensi larval habitats. Predicted capture point habitat risk surfaces were mapped using Python and analyzed with spatial autocorrelation metrics such as Local Moran’s I and Getis-Ord Gi*, which identified statistically significant clusters of high predicted stratified, An. stephensi habitat (hotspots) and low predicted risk (cold spots). Our assumption was that in large metropolitan areas such as downtown urban Kisumu, stratified clusters may correspond to patterns in breeding sites, transportation distribution networks, socioeconomic vulnerability, or service access gaps. The machine-learning predictions combined with spatial clustering produced a county-wide geospatial An. stephensi's habitat map is related to transmission-related risks concentrated geographically.
Afterwards, neighboring georeferenced, sub-county-level larval
An. stephensi habitat stratified, capture point, hot and cold spots were identified based on the stratified, measured, epidemiological and entomological sampled, non-time series, dependent discrete integer values, in an n-by-n binary geographic connectivity/weights matrix, C (Equation (
33)).
(33)
Thereafter, eigen-autocorrelation residual coefficients were expressible in terms of the product moment correlation coefficient formula. However, with neighboring capture point values
y replacing the value of each sampled variable
x (
33) became:
(34)
The left-hand expression converted to the right side in the stratified, eigen-autocorrelation, habitat, spectral signature, model, by substituting the numerator term only when a 1 appeared in the matrix C and by averaging the numerator cross-product terms over the total number of sampled estimator determinants, denoted in matrix C, i.e.
. The denominator of the revised expression of Equation (
33) was the sample variance of Y,
. Coupling this with part of the accompanying numerator term rendered
(35)
Standard normal deviate (i.e. z score) notation
was calculable. The right-hand expression of Equation (
33) was Moran’s coefficient. The coefficient resembled a Pearson product-moment correlation coefficient, ranging in value from −1 to +1. As indicated by −(
n − 1)/
n, Moran’s was influenced by a distribution of a number of georeferenced, stratified, neighbors (via
and via
. A standard bivariate regression included an intercept regression of
on Zy.
Figure 5. Potential Hot & Cold spots of Anopheles Stephensi habitats in Kenya.
A Box-Cox type of power transformation was subsequently employed for normal approximation analysis purposes, so that the frequency distributions of the georeferenced, sampled, capture point, larval
An. stephensi's habitat estimator determinants approximated a bell-shaped curve. The eigen-spatial filter construction methodology transformation procedure was used, as proposed by
| [15] | Jacob, B. G., Izureta, R., Bell, J., Parikh, J., Louh, D., Gates, T., Murray, K., White, L., and Aceng, J. R. Approximating Non-Asymptoticalness, Skew Heterr (2023).
https://doi.org/10.5923/j.ajms.20231301.01 |
[15]
, which depends on the eigenfunctions of matrix (I−11
T/n)C(I−11
T/n), where I denoted the identity matrix, 1 was an n-by-1 vector of ones, and T denoted matrix transpose, a term which appeared in the numerator of Moran’s Coefficient. According to
| [15] | Jacob, B. G., Izureta, R., Bell, J., Parikh, J., Louh, D., Gates, T., Murray, K., White, L., and Aceng, J. R. Approximating Non-Asymptoticalness, Skew Heterr (2023).
https://doi.org/10.5923/j.ajms.20231301.01 |
[15]
, the first eigenvector,
E1, is the set of numerical values with the largest spatial indices achievable by any set for the spatial arrangement defined by the connectivity matrix C. The second eigenvector is the set of values with the largest achievable autocorrelated values by any set that is uncorrelated with E
1. The third eigenvector is the third such set of values, and so on. This sequential construction of eigenvectors continued in the stratified, capture point,
An. stephensi habitat prognosticative model through E
n, where the set of measured stratified, eigen-autocorrelated,
An. stephensi, habitat estimator determinant, with the largest positive to negative capture point values, was rendered. Moran’s Coefficient was achieved by the set uncorrelated with the preceding (
n − 1) eigenvectors.
To identify the stratified geospatial larval An. stephensi habitat clusters that could be uncovered with eigen-spatial filtering, Thiessen polygon surface partitions were generated to construct geographic neighbor matrices in PySAL, whereby each hot and cold spot in Kisumu County was denoted by matrix C. Entries in matrix C were 1 if two An. stephensi habitat sampled, georeferenced, capture points shared a common Thiessen polygon boundary and 0 otherwise. The diagonal was coded 0. Next, the linkage structure for each surface was edited to remove unlikely geographic neighbors to identify pairs of georeferenceable capture point locations sharing a common Thiessen polygon boundary. Eigenvectors of a modified version of this matrix C furnished synthetic variates that constituted distinct habitat map patterns representing the full range of non-zero, Gaussian, eigen-autocorrelated possibilities. Attention was restricted to those map patterns associated with at least a minimum level of eigenized autocorrelation statistics, which, for implementation purposes here, was definable by Moran’s Coefficient using max > 0.25. Moran’s Coefficient denoted the jth value, and Moran’s Coefficient, maxx, the maximum value of the eigen-autocorrelation index. This threshold value allowed two candidate sets of eigenvectors to be considered for substantial positive and substantial negative eigen-autocorrelation, respectively.
We performed the Moran’s coefficient decomposition to quantitate whether or not equation (
33) can be replaced by:
(36)
This was determined by
K denoting some subset of the n eigenvectors derived from the eigen-spatial filter eigen-decomposition analyses of the remotely entomologically sampled, georeferenced, stratified,
capture point,
An. stephensi's habitat estimator determinants were chosen by the supervised selection criteria. The
was an n-by-
K matrix whose columns were the
K selected eigenvectors. This allowed dispensing with the
term by shifting spatial dependence effects to the large-scale variation term represented by E
i,Kβ, forcing
ρ to 0 in the eigen-autocorrelation capture point model. We let
αi be the constant
α. Only the stratified covariates were included in the habitat model specification. A link between Equations (
33) and (
34) was gleaned from the following algebraic manipulations:
where Λ was a diagonal matrix of eigenvalues whose order was the same as the corresponding eigenvectors in matrix E,
δ was a vector of coefficients where superscript T denoted the matrix transpose. A straightforward extension of Equation (
34) included the sub- county-level,
stratified,
An. stephensi habitat estimator determinant model specification; the version selected focused on the eigenvector eigen-spatial filter. The parameters
α and
β in Equation (
34) were subsequently estimated with the method of maximum likelihood.
The Moran scatter plot was also employed for the visualization of the zero/non-zero residual eigen-autocorrelation in the sub-county-level empirical, sampled georeferenced entomological datasets. The Moran's scatterplot was based upon Equation (
36), as was a plot of the pairs (
, linking it directly to standard correlation and regression analyses of the larval sampled
An. stephensi habitat estimator determinants. To construct the scatter plot, the count values of the stratified entomological data, Y, were converted to
z scores. Next, those adjacent or nearby georeferenced sub-county-level
z-score values of Y were summed using the matrix product Czy, where Zy was the vector concatenation of the individual Zy values. Finally, the coordinates of the sampled CC estimator determinants were quantified in eigenvector eigen-geospace. We noted that
),
i= 1, 2,
n, was plottable on the graph, where the vertical axis was Czy and the horizontal axis was
Zy. Significant hot spots (p < 0.01) were concentrated in Kisumu, characterized by higher socioeconomic deprivation and lower screening coverage. In contrast, statistically significant cold spots (p < 0.01) were observed in the entomological intervention site corresponding to areas with higher socioeconomic status and greater access to preventive services.
The vertical axis point was written in the stratified, capture point, larval An. stephensi's habitat model as , where resulting in an averaging of the surrounding nearby habitat sampled count values (yj − ȳ)/sY for each capture point i's using (yi − ȳ)/sY. The quantities were employed as the slope of the regression line for Moran's scatterplot, which in this experiment was the unstandardized Moran’s coefficient. The wij values were organized in an n-by-n matrix w, which was the row-standardized counterpart to matrix C. Finally, the coordinate pairs i= 1, 2, n, were plotted on the graph where the vertical axis was CZY and the horizontal axis was Zy. The unstandardized Moran’s coefficient was computed using the standard bivariate regression technique Czy on Zy, yielding bcy/y while specifying a non-intercept option. A no-intercept regression of C1 provided the discrete integer value for each, producing the standardization coefficient bc1/1− Moran coefficient =bcy/y/bc1/1.
Table 8. Global Moran’s I.
Statistic | Value |
Moran’s Index | -0.158504 |
Expected Index | -0.018519 |
Variance | 0.004646 |
z-score | -2.053665 |
p-value | 0.040008 |
Figure 6. Eigen-spatial Autocorrelation Report.