General considerations

Scope

This section primarily builds on established land-cover validation frameworks and good-practice recommendations for design-based inference [1], [2]. The central principle is that reference data used for validation should ideally be derived from a probability sampling design with known, non-zero inclusion probabilities, enabling statistically rigorous and unbiased accuracy assessment. While many of these concepts originated in land-cover validation, they remain broadly applicable across a wide range of geospatial products.

The discussion covers both categorical and continuous map products. Categorical products include multi-class land-cover maps as well as thematic single-class products such as forest extent or forest change maps. Continuous products include variables such as biomass and forest height. Although the general validation principles are shared, specific aspects of response design and accuracy assessment may differ across product types.

Not all mapping applications lend themselves to probability sampling, however. In cases where validation must rely on existing non-probability data sources (e.g., opportunistic field campaigns), alternative approaches are necessary. We discuss these in Section Validation Beyond Probability Sampling, along with recommended best practices for situations where design-based inference (based on probability sampling) is not feasible. ## Common sources of confusion Despite the central role of validation in establishing the scientific credibility of geospatial products, accuracy assessment methodology is frequently under-reported, and reference datasets are rarely made publicly available [3], [4]. Several recurring sources of confusion contribute to this problem. The first is the conflation of algorithm-derived uncertainty with map accuracy. Metrics produced by ensemble variance, posterior probabilities, or model confidence scores quantify internal model uncertainty, not agreement with ground truth. These measures reflect what the model “thinks” about its own predictions and should not be interpreted as accuracy estimates [5]. The second is the conflation of model evaluation with map validation. In many studies, standard ML practices (such as evaluating performance on a held-out test set) are presented as evidence of map quality. Model evaluation is indeed an important component of map development, as the final map is generated through spatial application of the model across a large geographic domain. Consequently, model performance and map quality are inherently related. However, they are not equivalent. Model evaluation primarily measures how well the model generalizes to a held-out test dataset, whereas map validation concerns the quality of the final spatial product across the full mapped domain. A held-out test set is typically not a probability sample of the map and may not represent the full range of landscapes, class distributions, and edge cases encountered in the final product. Even when the training sample is itself a probability sample and a portion is set aside for evaluation, this is generally not recommended: biases and errors present in the training data may transfer to the reference data [2]. More broadly, when map and reference data originate from the same source or interpretation process, their errors tend to be correlated, and correlated errors inflate quality estimates, while independent errors deflate them [6]. For these reasons, strong test-set performance does not guarantee, and cannot substitute for, an independent assessment of map quality. The third is the assumption that comparison against another existing map product constitutes validation. In principle, validation should rely on reference data of higher quality than the product being evaluated, such as field observations or carefully interpreted high-resolution imagery. Existing map products, however, are often themselves the result of inference procedures, including machine learning models or rule-based gridded approaches, and therefore contain their own uncertainties and biases. While map-to-map comparisons can provide valuable qualitative insights and help identify systematic differences between products [7], they should not be interpreted as ground truth validation. What these comparisons measure is inter-product consistency, not map quality. They are more accurately described as consistency checks and should be reported as such.

[1]
P. Olofsson, G. M. Foody, M. Herold, S. V. Stehman, C. E. Woodcock, and M. A. Wulder, Good practices for estimating area and assessing accuracy of land change,” Remote Sensing of Environment, vol. 148, pp. 42–57, May 2014, doi: 10.1016/j.rse.2014.02.015.
[2]
A. Tyukavina, S. V. Stehman, A. H. Pickens, P. Potapov, and M. C. Hansen, Practical global sampling methods for estimating area and map accuracy of land cover and change,” Remote Sensing of Environment, vol. 324, p. 114714, 2025.
[3]
S. V. Stehman and G. M. Foody, Key issues in rigorous accuracy assessment of land cover products,” Remote Sensing of Environment, vol. 231, p. 111199, 2019, doi: 10.1016/j.rse.2019.05.018.
[4]
G. M. Foody, Status of land cover classification accuracy assessment,” Remote Sensing of Environment, vol. 80, no. 1, pp. 185–201, Apr. 2002, doi: 10.1016/s0034-4257(01)00295-4.
[5]
A. Tyukavina et al., Land Cover and Change Map Accuracy Assessment and Area Estimation Good Practices Protocol. 2025. doi: 10.5067/doc/ceoswgcv/lpv/lc.001.
[6]
G. M. Foody, Assessing the accuracy of land cover change with imperfect ground reference data,” Remote Sensing of Environment, vol. 114, no. 10, pp. 2271–2285, Oct. 2010, doi: 10.1016/j.rse.2010.05.003.
[7]
A. H. Strahler et al., Global Land Cover Validation: Recommendations for Evaluation and Accuracy Assessment of Global Land Cover Maps,” European Commission, Joint Research Centre, Institute for Environment; Sustainability, Ispra, Italy, GOFC-GOLD Report No. 25, 2006.