Ongoing research

Spatial analysis and machine learning for river water quality

A developing multi-river Bangladesh research framework that converts published heavy-metal observations into a structured spatiotemporal database, links sampling locations to environmental predictors, and prepares the data foundation for machine-learning analysis.

Supervisor: Prof. Md. Saiful Islam · KUETLiterature synthesis completedGeospatial predictor extraction in progress

Research direction

Heavy-metal measurements in Bangladeshi rivers are distributed across individual studies, locations, years, sample types, analytical methods, and reporting conventions. That fragmentation makes it difficult to ask broader spatial questions about how environmental conditions may be associated with observed contamination.

This project is building the data infrastructure needed to investigate those relationships. The current objective is to integrate published water and sediment heavy-metal observations with climatic, land-use, population, and other spatial predictors, then use machine-learning methods to explore and predict concentration patterns.

Current research question: can a harmonized literature-derived dataset, augmented with multi-scale geospatial predictors, explain meaningful variation in river heavy-metal concentrations across Bangladesh?

Phase 1: literature review → structured database

The literature-review stage is already substantially complete. Rather than keeping the review as narrative notes, I extracted the reported information into a structured workbook so that each sampling record can later be joined to spatial predictors and used in analysis.

The working database currently contains 557 populated records from 28 DOI-tagged papers, covering 12 rivers and 54 fields. Published studies represented in the workbook extend through 2025, while the reported observation-year field currently spans 2010–2022 where study years are available.

557extracted study records
28DOI-tagged source papers
12Bangladesh rivers represented
54database fields
Why this matters: the literature review is being treated as a reproducible data-building step, not only as background reading. The output becomes the observational layer for the later geospatial and machine-learning workflow.

What is being extracted

The workbook combines bibliographic, spatial, temporal, analytical and environmental information. Missing values are preserved rather than silently filled, allowing later preprocessing decisions to remain explicit.

Paper metadataPaper ID, DOI, journal, author & year
Spatial contextRiver, division, station, latitude, longitude, population field
Time & sampleObservation year, season, water/sediment, analytical method
Physicochemical datapH, EC, DO, BOD, COD, salinity, texture, OM and other reported variables
Heavy metalsAs, Cr, Cd, Pb, Cu, Zn, Ni, Fe, Mn and Co
Source notesSampling-context notes and reporting details retained for traceability
Sample coverageBoth sediment and water observations are represented
Data provenancePaper and DOI fields support record-level trace-back

Representative database snapshot

SourceRiverLatitudeLongitudeYearSampleAsCrCdPb
Islam et al., 2022Old Brahmaputra24°46′11.37″ N90°24′00.29″ E2021Sediment1.0730.514.3632.29
Islam et al., 2022Old Brahmaputra24°44′58.46″ N90°25′24.89″ E2021Sediment1.9337.284.9226.04
Nargis et al., 2019Buriganga23°44′36.61″ N90°20′45.08″ E2015Sediment0.2876.440.233.93
Nargis et al., 2019Buriganga23°42′27.55″ N90°24′13.07″ E2015Sediment0.2242.010.2440.87
Hasan et al., 2024Karnaphuli22.521519392.034745012022Water—0.0710.0010.09
Hasan et al., 2024Padma24.505672188.875643212022Water—0.03BDL0.197
Hasan et al., 2024Meghna23.632911790.518905522022Water—0.040.0810.271
Hasan et al., 2024Shitalakhya23.609978590.615282592022Water—0.090.010.235
Ali et al., 2016Karnaphuli22.5354472292.406627782014Sediment13.1757.311.1035.25
Ali et al., 2016Karnaphuli22.31513291.8194242014Sediment13.38111.481.4061.86
Islam et al., 2015Paira22°27′03.88″ N90°26′50.11″ E2012Sediment5.4460.7915
Islam et al., 2015Paira22°24′42.53″ N90°26′37.78″ E2012Sediment17571.144
Proshad et al., 2021Rupsa22.80 N89.54 E2017Water0.006050.008870.001380.00732

The snapshot uses records with reported latitude and longitude from the working literature-extraction workbook. Values are shown as recorded in the source database; concentration units depend on sample context (water or sediment) and are intentionally not harmonized in this preview.

Research progress

PHASE 01

Literature review & record extraction

Completed

Relevant studies were reviewed and their source metadata, sampling context, water/sediment observations, analytical methods, physicochemical variables and heavy-metal concentrations were extracted into the working database.

PHASE 02

Geospatial preparation & environmental predictor extraction

In progress

Current work is preparing reported sampling locations for spatial analysis and extracting environmental information around them, including rainfall, population and buffer-zone data at 200 m, 500 m and 1,000 m radii. Land-use and other spatial predictors are part of the broader project design described in the research brief.

PHASE 03

Data harmonization & modeling dataset preparation

Required before modeling

The literature-derived records use different reporting conventions, sample types and units. These differences will need to be handled explicitly before a common modeling dataset is finalized; missing values will remain traceable to their sources.

PHASE 04

Machine-learning model development

Planned

After predictor assembly and data-quality checks, the next phase is to design and evaluate machine-learning models for heavy-metal concentration prediction and investigation of environmental drivers.

Multi-scale environmental predictor extraction

The current geospatial stage is designed around three buffer scales—200 m, 500 m and 1,000 m—around sampling locations. The purpose is not to assume that one radius is universally correct, but to test whether local versus broader landscape context changes the relationship between environmental predictors and observed concentrations.

200 m

Working buffer radius

Environmental variables are being extracted within a 200 m radius around reported sampling locations.

500 m

Working buffer radius

The same predictor families are being assembled at a 500 m radius for multi-scale comparison.

1,000 m

Working buffer radius

A 1,000 m radius is also being extracted so the final modeling stage can evaluate which spatial scale is most useful.

Predictor families being assembled

RainfallClimatic context aligned to spatial/temporal observations where feasible
PopulationHuman-pressure proxy around sampled river locations
Land use / land coverUrban, agricultural and other surrounding surface characteristics
Spatial contextRiver, sampling location and basin-scale descriptors

The machine-learning phase is next

The supplied research plan places machine-learning model design after completion of the literature-derived dataset and environmental predictor extraction. The goal is to predict heavy-metal concentrations using basin- and site-scale environmental predictors.

A specific algorithm, final feature set, train/test strategy and performance metric have not yet been finalized in the supplied project materials. They are therefore not presented here as completed methodological decisions.

Published studies ↓ Structured heavy-metal database ↓ Location preparation + data harmonization ↓ 200 m / 500 m / 1000 m environmental predictors ↓ Final modeling dataset ↓ Machine-learning model design ↓ Prediction and interpretation

Data-quality and comparability challenges

A literature-derived environmental dataset is inherently heterogeneous. The workbook already shows differences in capitalization, division spelling, sample descriptions, analytical-method reporting, missing coordinates, below-detection-limit notation, observation-year availability, and the set of metals/physicochemical variables reported by each paper.

Those differences are not treated as information to hide. They have to be resolved or explicitly retained during preparation of the modeling dataset. The working plan includes category/name normalization, unit auditing, explicit handling of BDL and missing values, and source-level traceability so each extracted value can be checked against the paper from which it came.

Still to be decided: the final preprocessing rules, model family and validation strategy are part of the next research stage and are not claimed as completed work on this page.

Next milestones

  • Complete buffer-based rainfall, population and land-use extraction.
  • Finish metadata/category harmonization and location-quality checks.
  • Define metal-specific modeling datasets based on available coverage and comparability.
  • Design the machine-learning modeling and validation workflow.
  • Evaluate model performance and interpret the environmental predictors used in the final models.
  • Develop spatial prediction outputs after the modeling framework is validated.

This page will be updated as the project moves from data construction into validated modeling; planned outputs are deliberately separated from completed work.