Is your feature request related to a problem? Please describe.
Grid models built from "latest" OSM data cannot be reproduced. OSM changes continuously, so the same config run a month apart yields a different network. Results, validation and published studies cannot be re-run or compared against a fixed input. Pinning the OSM snapshot date is a prerequisite for reproducible runs.
Describe the solution you'd like
A config-selectable OSM data source with an explicit snapshot date, and snapshots stored per date so runs with different dates never overwrite each other.
Proposed config (same shape as the one already implemented in PyPSA-Earth):
osm_data:
source: historical # latest | historical | custom
target_date: "2020-01-01" # YYYY-MM-DD, required for historical
latest: current OSM data (default, unchanged behaviour)
historical: data as of target_date
custom: user-provided .pbf and optional pre-filtered power files
Raw data is stored per snapshot, e.g. data/osm/<subdir>/ where <subdir> is YYYYMM, latest or custom.
Acceptance criteria:
Prior work (what is already done and can be reused)
pypsa-earth-osm
- #5 historical OSM data integration (from earth-osm#63)
- #16 separate OSM data into date-specific directories
- #27 test config for historical data (
test/config.historical.yaml)
pypsa-earth
- #2005 ports the feature with the
osm_data config section above; requires earth-osm>=3.0.2 (open)
Describe alternatives you've considered
- Pinning a downloaded
.pbf by hand (custom source): works, but the snapshot date is not recorded in the config and is easy to lose.
- Committing/archiving the raw data per study: heavy, and does not scale across users.
Open questions
- Which backend should serve historic data here (earth-osm historical feature vs Geofabrik history files vs Overpass
date: queries)? Each has a different earliest date and coverage.
- Should the resolved snapshot date be written into the output metadata so results are traceable?
- Granularity of the snapshot: month (
YYYYMM, as in PyPSA-Earth) or exact date?
Additional context
Related: retrieve.target_date in grid-builder (workflow/scripts/_schema.py) currently exists but is only honoured for retrieve.source: geofabrik, so the Overpass path is not reproducible today.
Is your feature request related to a problem? Please describe.
Grid models built from "latest" OSM data cannot be reproduced. OSM changes continuously, so the same config run a month apart yields a different network. Results, validation and published studies cannot be re-run or compared against a fixed input. Pinning the OSM snapshot date is a prerequisite for reproducible runs.
Describe the solution you'd like
A config-selectable OSM data source with an explicit snapshot date, and snapshots stored per date so runs with different dates never overwrite each other.
Proposed config (same shape as the one already implemented in PyPSA-Earth):
latest: current OSM data (default, unchanged behaviour)historical: data as oftarget_datecustom: user-provided.pbfand optional pre-filtered power filesRaw data is stored per snapshot, e.g.
data/osm/<subdir>/where<subdir>isYYYYMM,latestorcustom.Acceptance criteria:
target_dategives identical raw OSM input across runstarget_dateis validated (format, not in the future, not before the earliest supported snapshot)target_date: 2020-01-01) runs in CIPrior work (what is already done and can be reused)
pypsa-earth-osmtest/config.historical.yaml)pypsa-earthosm_dataconfig section above; requiresearth-osm>=3.0.2(open)Describe alternatives you've considered
.pbfby hand (customsource): works, but the snapshot date is not recorded in the config and is easy to lose.Open questions
date:queries)? Each has a different earliest date and coverage.YYYYMM, as in PyPSA-Earth) or exact date?Additional context
Related:
retrieve.target_dateingrid-builder(workflow/scripts/_schema.py) currently exists but is only honoured forretrieve.source: geofabrik, so the Overpass path is not reproducible today.