Skip to content

Support pinned historic OSM snapshots for reproducibility #10

Description

@GbotemiB

Is your feature request related to a problem? Please describe.

Grid models built from "latest" OSM data cannot be reproduced. OSM changes continuously, so the same config run a month apart yields a different network. Results, validation and published studies cannot be re-run or compared against a fixed input. Pinning the OSM snapshot date is a prerequisite for reproducible runs.

Describe the solution you'd like

A config-selectable OSM data source with an explicit snapshot date, and snapshots stored per date so runs with different dates never overwrite each other.

Proposed config (same shape as the one already implemented in PyPSA-Earth):

osm_data:
  source: historical      # latest | historical | custom
  target_date: "2020-01-01"   # YYYY-MM-DD, required for historical
  • latest: current OSM data (default, unchanged behaviour)
  • historical: data as of target_date
  • custom: user-provided .pbf and optional pre-filtered power files

Raw data is stored per snapshot, e.g. data/osm/<subdir>/ where <subdir> is YYYYMM, latest or custom.

Acceptance criteria:

  • Same config + same target_date gives identical raw OSM input across runs
  • Different dates are stored in separate directories and do not overwrite each other
  • target_date is validated (format, not in the future, not before the earliest supported snapshot)
  • A test config exercising the historical path (e.g. target_date: 2020-01-01) runs in CI
  • Docs describe the options and the reproducibility guarantee and its limits

Prior work (what is already done and can be reused)

  • pypsa-earth-osm
    • #5 historical OSM data integration (from earth-osm#63)
    • #16 separate OSM data into date-specific directories
    • #27 test config for historical data (test/config.historical.yaml)
  • pypsa-earth
    • #2005 ports the feature with the osm_data config section above; requires earth-osm>=3.0.2 (open)

Describe alternatives you've considered

  • Pinning a downloaded .pbf by hand (custom source): works, but the snapshot date is not recorded in the config and is easy to lose.
  • Committing/archiving the raw data per study: heavy, and does not scale across users.

Open questions

  • Which backend should serve historic data here (earth-osm historical feature vs Geofabrik history files vs Overpass date: queries)? Each has a different earliest date and coverage.
  • Should the resolved snapshot date be written into the output metadata so results are traceable?
  • Granularity of the snapshot: month (YYYYMM, as in PyPSA-Earth) or exact date?

Additional context

Related: retrieve.target_date in grid-builder (workflow/scripts/_schema.py) currently exists but is only honoured for retrieve.source: geofabrik, so the Overpass path is not reproducible today.

Activity

  1. ekatef commented on Oct 5, 2026

    @ekatef
    Collaborator

    Hi @GbotemiB thanks a lot for the analysis! Totally support the need to have the historical feature. As you have shown in the GridImpact studies, the effect can be quite pronounced. Don't you have a couple of pictures to showcase how this feature can actually work?

  2. GbotemiB commented on Oct 5, 2026

    @GbotemiB
    Author

    Colombia (case-study example)

    The following images are grid connection for Colombia for 2020, 2024 and late 2025 showing improvement. With the historical-osm feature, the following maps can be reproduced, with applicability to other part of the world.

    Image

    Here is the image for 2020

    Image

    Here is the map for 2024

    Image

    Here is the map for late 2025

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions