Skip to content

Repository files navigation

OmniWaterMask

image image image Conda Recipe

OmniWaterMask is a Python library for high accuracy water segmentation in high to moderate resolution satellite imagery, supporting a wide range of resolutions, sensors, and processing levels.

Check out the paper here

Model version 2 is in beta

Version 2 is a new model, trained on seven datasets rather than two. It is the default from 0.7.0b1, which is a prerelease, so a normal install still gives you version 1. To try it:

uv add omniwatermask==0.7.0b1

The paper describes version 1. To reproduce the published model, pass model_version=1 and install the legacy extra, which is described below.

Features

  • Process imagery resolutions from 0.2 m to 50 m.
  • Any imagery processing level
  • Only requires Red, Green, Blue and NIR bands
  • Known to work well with Sentinel-2, Landsat 8, PlanetScope, Maxar and NAIP

Try in Colab

Colab_Button

How it works

OmniWaterMask combines a sensor agnostic deep learning segmentation model with NDWI and vector data to find water in satellite and aerial imagery.

Installation

Install it with an environment manager such as conda or uv, to keep it clear of your other packages.

Install the package using pip

pip install omniwatermask

Install the package using uv

uv add omniwatermask

Create a new conda environment and install from conda-forge

conda create -n owm python=3.12
conda activate owm
conda install -c conda-forge omniwatermask

Install the package from source

pip install git+https://github.com/DPIRD-DMA/OmniWaterMask.git

Running model versions below 2 (optional)

Model versions below 2 are fastai models, and fastai is not installed by default. The current model does not need it, so install this only to run an older version. These are alternatives. Use whichever matches how you installed OmniWaterMask.

With pip:

pip install "omniwatermask[legacy]"

With uv:

uv add omniwatermask --extra legacy

With conda:

conda install conda-forge::omniwatermask conda-forge::fastai

Usage

Pass a list of geotiff files to make_water_mask with the band order for the Red, Green, Blue and NIR bands. It writes each prediction to disk as a geotiff alongside its input, and returns the list of prediction paths.

from pathlib import Path
from omniwatermask import make_water_mask

scene_paths = [Path("path/to/scene1.tif"), Path("path/to/scene2.tif")]

# Predict water masks for scenes
water_mask_path = make_water_mask(
    scene_paths=scene_paths,  # you can pass a list of images
    band_order=[1, 2, 3, 4],  # band order of the input images, expects RGB+NIR
)

Output

Output classes:

  • 0 = non-water
  • 1 = water

Examples

Example notebooks are available in the examples/ directory:

Usage tips

  • OWM needs an internet connection, because it downloads vector data.
  • Which Overture release to read is resolved once and reused for the rest of the process, from Overture's release catalogue when it is reachable and otherwise from the newest release in Overture's S3 bucket that carries every theme OWM reads. It is rediscovered if a fetch later fails against it, so a long-running process picks up a new release after Overture prunes the old one. Releases are not pinned. Overture retains roughly two releases (~60 days) and prunes the rest, so a hardcoded release stops resolving within a couple of months. This also means an old run cannot be reproduced by pinning a release. The local vector cache is what makes a target set reproducible.
  • Vector data comes from Overture Maps by default. If you would rather query OpenStreetMap live through the Overpass API, set vector_source="osm". Overture serves static monthly GeoParquet releases from cloud storage, so it avoids the rate limits and timeouts Overpass returns on large or dense bounding boxes. The underlying data is largely the same, since Overture's water and road layers are derived from OSM, though its building footprints add machine-learning-derived data beyond OSM. Note that Overture files a few landforms (cape, blowhole, shoal) under its water theme; OWM filters these out so they are not treated as water.
  • If a scene's vector data cannot be fetched, that scene is skipped rather than processed without it. A mask built without its vector targets looks plausible but is quietly worse. Overture fetches retry transient failures first (3 attempts, 2s then 4s apart); failures that will not improve on a retry, such as being unable to determine which Overture release to read, are raised immediately instead of consuming the backoff. A skipped scene is logged at ERROR, is left out of the returned list of output paths, and has no file written, so re-running the same call reprocesses it while the rest of the batch is untouched. Because the two sources are largely interchangeable, an outage in one is worth trying the other for, and the error messages say so.
  • Use hardware acceleration if you have it:
    • NVIDIA GPU
    • Apple Silicon Mac
    • Other PyTorch-compatible accelerators
  • inference_dtype defaults to "auto", which times the inference device once and drops to a reduced-precision dtype only where that is measurably faster. Pass an explicit dtype (such as "bf16" or torch.float32) to override the measurement.
  • If you run out of VRAM even at batch_size=1, set mosaic_device to "cpu".
  • Improve accuracy by passing known water body locations to aux_vector_sources, as a list of paths to your water polygon datasets.
  • Reduce false positives by passing vector data for things often mistaken for water, such as buildings and roads, to aux_negative_vector_sources.
  • For scenes with no-data regions, set no_data_value so those pixels are read as no-data rather than as imagery.

Cloudy imagery

If you are working with cloudy imagery, either:

  • use a temporal mosaic that is already cloud and cloud-shadow free (e.g. via s2mosaic for Sentinel-2), or
  • apply a high quality cloud and cloud shadow mask and set those pixels to 0 (the no_data_value) before running OWM.

This matters because OWM optimises its detection thresholds both locally (per region or patch) and globally (across the whole scene). Cloud and cloud-shadow pixels are out-of-distribution and can skew those optimisations, so bad data in one part of a scene can degrade the water prediction in other, otherwise-clean parts. Masking those pixels to no-data removes them from the optimisation entirely.

OmniCloudMask is a good choice for the masking step. See the cloudy Sentinel-2 example for an end-to-end mask-then-infer workflow.

Parameters

  • scene_paths: List of paths or single path (supports both Path and string types) to the input satellite/aerial imagery

  • band_order: List of integers specifying the band order for input imagery (e.g., [1,2,3,4] if your input image is stored with band order red, green, blue then NIR data). This tells OWM which bands correspond to Red, Green, Blue, and Near-Infrared channels

  • batch_size: Number of patches processed simultaneously during inference. Default is 1, increase for better GPU utilization

  • version: Version identifier for the output files. Defaults to current OmniWaterMask version

  • output_dir: Optional path for output files. If not specified, outputs are saved alongside input files

  • mosaic_device: Device for mosaic operations ("cpu", "cuda" or "mps"). Defaults to system's default device

  • inference_device: Device for model inference ("cpu", "cuda" or "mps"). Defaults to system's default device

  • aux_vector_sources: List of paths to supplementary water body vector data to aid detection

  • aux_negative_vector_sources: List of paths to vector data marking areas commonly misidentified as water

  • inference_dtype: Data type for inference operations. Defaults to "auto", which measures the inference device and picks the fastest dtype that is not slower than float32; pass a dtype or dtype string to set it explicitly

  • no_data_value: Value indicating no-data regions in the input imagery. Defaults to 0

  • inference_patch_size: Size of image patches for inference. Defaults to 1000 pixels

  • inference_overlap_size: Overlap between adjacent patches during inference. Defaults to 300 pixels

  • overwrite: Whether to overwrite existing output files. Defaults to True

  • use_cache: Whether to cache vector data processing results. Defaults to True

  • use_osm_building: Whether to use building data to reduce false positives. Defaults to True

  • use_osm_roads: Whether to use road data to reduce false positives. Defaults to True

  • vector_source: Where water, road and building vectors come from, either "overture" (Overture Maps GeoParquet) or "osm" (OpenStreetMap via the Overpass API). Defaults to "overture"

  • include_ocean: Whether Overture ocean polygons count as positive water targets. These cover everything seaward of the OSM coastline, which the OSM tag set does not provide. Set to False if coastline/tide offsets cause false positives on your scenes. Only applies when vector_source="overture". Defaults to True

  • cache_dir: Directory for storing cached vector data. Defaults to "OWM_cache" in current directory

  • destination_model_dir: Directory to save the model weights. Defaults to None

  • model_download_source: Source from which to download the model weights. Defaults to "hugging_face", can also be "google_drive".

  • model_version: Which published model version to use. Defaults to the newest in the packaged index. Versions below 2 are fastai models and need the legacy extra installed (see above).

Cache maintenance

Cached vectors are stored in a generation, a database and a parquet directory whose names carry a version. A release that changes how vectors are stored bumps the generation, so older entries are ignored rather than migrated, and the old files stay on disk: a full copy of the cache per bump.

prune_stale_cache(cache_dir) reclaims that space. It deletes generations below the current one, along with any parquet in the current generation that no entry points at, and returns the number of files deleted. Generations above the current one are left alone. They belong to a newer install sharing the directory, so they are in use rather than obsolete.

from omniwatermask import prune_stale_cache

deleted = prune_stale_cache("OWM_cache")

It is deliberately manual rather than automatic: those files are the only copy an older install would still read, so pruning means a downgrade refetches.

Changelog

See CHANGELOG.md for a full list of changes across versions.

Contributing

Contributions are welcome! Please submit a pull request or open an issue to discuss any changes.

Development setup

Clone the repository and install the dependencies (including the dev group) with uv:

uv sync --all-extras --dev

Optionally install the git hooks (ruff lint/format on commit, mypy + the fast tests on push):

uv run pre-commit install
uv run pre-commit install --hook-type pre-push

Running the tests

Tests use pytest. The fast suite (unit tests + model-mocked pipeline tests) runs in a few seconds and is what CI runs by default:

uv run pytest                              # full fast suite
uv run pytest tests/test_orchestration.py  # one file
uv run pytest -k make_water_mask           # match by name

End-to-end tests that download the real model weights and run inference on real imagery are marked e2e and excluded by default (see addopts in pyproject.toml). To run them explicitly:

uv run pytest -m e2e                        # only the e2e/inference tests
uv run pytest -m ""                         # everything, including e2e

These download every published model version from both download sources and run real inference, so allow around ten minutes and a few hundred MB. Sync with --all-extras before running them, or the model version 1 cases skip for want of fastai rather than failing.

Lint, format and type-check:

uv run ruff check .
uv run ruff format .
uv run mypy omniwatermask/

For maintainers, pushing a version tag (e.g. git tag v0.4.4 && git push --tags) builds the package and publishes it to PyPI via GitHub Actions trusted publishing. No tokens are required.

License

This project is licensed under the MIT License

Acknowledgements

Special thanks to the authors of the datasets OmniWaterMask is trained on.

Model version 2 is trained on seven:

Model version 1, the model described in the paper, is trained on S1S2-Water and FLAIR #1.

About

Python library for high-accuracy water segmentation in satellite and aerial imagery, combining deep learning with NDWI and vector data for robust detection across multiple sensors and resolutions.

Topics

Resources

Stars

48 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages