Cross-validation for geospatial data: Estimating generalization performance in geostatistical problems

Jing Wang; Laurel Hopkins; Tyler Hallman; W Douglas Robinson; Rebecca Hutchinson

Cross-validation for geospatial data: Estimating generalization performance in geostatistical problems

Research output: Contribution to journal › Article › peer-review

Electronic versions

Documents

1149_cross_validation_for_geospatia
Accepted author manuscript, 4.07 MB, PDF document

Jing Wang
Oregon State University
Laurel Hopkins
Oregon State University
Tyler Hallman
School of Environmental & Natural Sciences
W Douglas Robinson
Oregon State University
Rebecca Hutchinson
Oregon State University

Geostatistical learning problems are frequently characterized by spatial autocorrelation in the input features and/or the potential for covariate shift at test time. These realities violate the classical assumption of independent, identically distributed data, upon which most cross-validation algorithms rely in order to estimate the generalization performance of a model. In this paper, we present a theoretical criterion for unbiased cross-validation estimators in the geospatial setting. We also introduce a new cross-validation algorithm to
evaluate models, inspired by the challenges of geospatial problems. We apply a framework for categorizing problems into different types of geospatial scenarios to help practitioners select an appropriate cross-validation strategy. Our empirical analyses compare cross-validation algorithms on both simulated and several real datasets to develop recommendations for a variety of geospatial settings. This paper aims to draw attention to some challenges that arise in model evaluation for geospatial problems and to provide guidance for users.

Original language	English
Journal	Transactions on Machine Learning Research
Publication status	Published - 4 Oct 2023

Total downloads

No data available

View graph of relations

Research Portal

Cross-validation for geospatial data: Estimating generalization performance in geostatistical problems

Electronic versions

Documents

Total downloads