Sensor data can look exact because it comes with timestamps and decimal places. That does not mean every reading is believable.
This R project checks campus weather-station data before using it for analysis. It combines the measurements with sensor state, timing, and physical rules.
- Missing and irregular timestamps
- Precipitation recorded when temperature suggests frozen conditions
- Precipitation while the sensor is tilted by more than two degrees
- Low battery voltage
- Air-temperature and solar-radiation values outside physical bounds
- Expected 15-minute and hourly sampling intervals
- An eight-observation moving average
- March–April precipitation after screening
Activity4.R
activity04/campus_weather.csv
activity04/meter_weather_metadata.csv
activity04/Sensor log.csv
This is a simple example of a rule I use across the rest of my work: check how the data were produced before modeling them. A bad sensor value should not quietly become a clean-looking chart or regression input.
The thresholds here are hand-built rules, not a machine-learning anomaly detector. A stronger pipeline would preserve every raw value, version the rules, attach a reason code to each flag, and measure how screening changes the final result.
install.packages(c("dplyr", "ggplot2", "lubridate"))The original script uses Posit Cloud paths and refers to an hourly soil-data folder that is not in this repo. Update the paths and skip that soil section if you are running only the weather-station analysis.