Research direction
Heavy-metal measurements in Bangladeshi rivers are distributed across individual studies, locations, years, sample types, analytical methods, and reporting conventions. That fragmentation makes it difficult to ask broader spatial questions about how environmental conditions may be associated with observed contamination.
This project is building the data infrastructure needed to investigate those relationships. The current objective is to integrate published water and sediment heavy-metal observations with climatic, land-use, population, and other spatial predictors, then use machine-learning methods to explore and predict concentration patterns.
Phase 1: literature review → structured database
The literature-review stage is already substantially complete. Rather than keeping the review as narrative notes, I extracted the reported information into a structured workbook so that each sampling record can later be joined to spatial predictors and used in analysis.
The working database currently contains 557 populated records from 28 DOI-tagged papers, covering 12 rivers and 54 fields. Published studies represented in the workbook extend through 2025, while the reported observation-year field currently spans 2010–2022 where study years are available.
What is being extracted
The workbook combines bibliographic, spatial, temporal, analytical and environmental information. Missing values are preserved rather than silently filled, allowing later preprocessing decisions to remain explicit.
Representative database snapshot
| Source | River | Latitude | Longitude | Year | Sample | As | Cr | Cd | Pb |
|---|---|---|---|---|---|---|---|---|---|
| Islam et al., 2022 | Old Brahmaputra | 24°46′11.37″ N | 90°24′00.29″ E | 2021 | Sediment | 1.07 | 30.51 | 4.36 | 32.29 |
| Islam et al., 2022 | Old Brahmaputra | 24°44′58.46″ N | 90°25′24.89″ E | 2021 | Sediment | 1.93 | 37.28 | 4.92 | 26.04 |
| Nargis et al., 2019 | Buriganga | 23°44′36.61″ N | 90°20′45.08″ E | 2015 | Sediment | 0.28 | 76.44 | 0.23 | 3.93 |
| Nargis et al., 2019 | Buriganga | 23°42′27.55″ N | 90°24′13.07″ E | 2015 | Sediment | 0.22 | 42.01 | 0.24 | 40.87 |
| Hasan et al., 2024 | Karnaphuli | 22.5215193 | 92.03474501 | 2022 | Water | — | 0.071 | 0.001 | 0.09 |
| Hasan et al., 2024 | Padma | 24.5056721 | 88.87564321 | 2022 | Water | — | 0.03 | BDL | 0.197 |
| Hasan et al., 2024 | Meghna | 23.6329117 | 90.51890552 | 2022 | Water | — | 0.04 | 0.081 | 0.271 |
| Hasan et al., 2024 | Shitalakhya | 23.6099785 | 90.61528259 | 2022 | Water | — | 0.09 | 0.01 | 0.235 |
| Ali et al., 2016 | Karnaphuli | 22.53544722 | 92.40662778 | 2014 | Sediment | 13.17 | 57.31 | 1.10 | 35.25 |
| Ali et al., 2016 | Karnaphuli | 22.315132 | 91.819424 | 2014 | Sediment | 13.38 | 111.48 | 1.40 | 61.86 |
| Islam et al., 2015 | Paira | 22°27′03.88″ N | 90°26′50.11″ E | 2012 | Sediment | 5.4 | 46 | 0.79 | 15 |
| Islam et al., 2015 | Paira | 22°24′42.53″ N | 90°26′37.78″ E | 2012 | Sediment | 17 | 57 | 1.1 | 44 |
| Proshad et al., 2021 | Rupsa | 22.80 N | 89.54 E | 2017 | Water | 0.00605 | 0.00887 | 0.00138 | 0.00732 |
The snapshot uses records with reported latitude and longitude from the working literature-extraction workbook. Values are shown as recorded in the source database; concentration units depend on sample context (water or sediment) and are intentionally not harmonized in this preview.
Research progress
Literature review & record extraction
CompletedRelevant studies were reviewed and their source metadata, sampling context, water/sediment observations, analytical methods, physicochemical variables and heavy-metal concentrations were extracted into the working database.
Geospatial preparation & environmental predictor extraction
In progressCurrent work is preparing reported sampling locations for spatial analysis and extracting environmental information around them, including rainfall, population and buffer-zone data at 200 m, 500 m and 1,000 m radii. Land-use and other spatial predictors are part of the broader project design described in the research brief.
Data harmonization & modeling dataset preparation
Required before modelingThe literature-derived records use different reporting conventions, sample types and units. These differences will need to be handled explicitly before a common modeling dataset is finalized; missing values will remain traceable to their sources.
Machine-learning model development
PlannedAfter predictor assembly and data-quality checks, the next phase is to design and evaluate machine-learning models for heavy-metal concentration prediction and investigation of environmental drivers.
Multi-scale environmental predictor extraction
The current geospatial stage is designed around three buffer scales—200 m, 500 m and 1,000 m—around sampling locations. The purpose is not to assume that one radius is universally correct, but to test whether local versus broader landscape context changes the relationship between environmental predictors and observed concentrations.
Working buffer radius
Environmental variables are being extracted within a 200 m radius around reported sampling locations.
Working buffer radius
The same predictor families are being assembled at a 500 m radius for multi-scale comparison.
Working buffer radius
A 1,000 m radius is also being extracted so the final modeling stage can evaluate which spatial scale is most useful.
Predictor families being assembled
The machine-learning phase is next
The supplied research plan places machine-learning model design after completion of the literature-derived dataset and environmental predictor extraction. The goal is to predict heavy-metal concentrations using basin- and site-scale environmental predictors.
A specific algorithm, final feature set, train/test strategy and performance metric have not yet been finalized in the supplied project materials. They are therefore not presented here as completed methodological decisions.
Data-quality and comparability challenges
A literature-derived environmental dataset is inherently heterogeneous. The workbook already shows differences in capitalization, division spelling, sample descriptions, analytical-method reporting, missing coordinates, below-detection-limit notation, observation-year availability, and the set of metals/physicochemical variables reported by each paper.
Those differences are not treated as information to hide. They have to be resolved or explicitly retained during preparation of the modeling dataset. The working plan includes category/name normalization, unit auditing, explicit handling of BDL and missing values, and source-level traceability so each extracted value can be checked against the paper from which it came.
Next milestones
- Complete buffer-based rainfall, population and land-use extraction.
- Finish metadata/category harmonization and location-quality checks.
- Define metal-specific modeling datasets based on available coverage and comparability.
- Design the machine-learning modeling and validation workflow.
- Evaluate model performance and interpret the environmental predictors used in the final models.
- Develop spatial prediction outputs after the modeling framework is validated.
This page will be updated as the project moves from data construction into validated modeling; planned outputs are deliberately separated from completed work.