Validation beyond probability sampling

The preceding sections describe the recommended framework for statistically rigorous map validation based on design-based inference and probability sampling. However, not all mapping applications lend themselves to this approach. For certain product types, particularly continuous-field estimates such as canopy height or above-ground biomass density, the cost and complexity of establishing a global probability sample with adequate reference quality may be prohibitive. In such cases, alternative validation strategies are necessary, though they cannot provide the same statistical guarantees as design-based inference. An alternative statistical framework is model-based inference. Where design-based inference derives its validity from the randomness of the sampling process, model-based inference relies on a statistical model that describes how map errors relate to predictors such as landscape characteristics or data availability. This allows it to work with non-probability reference data, such as opportunistic field plots or spatially sparse inventories, by extrapolating accuracy estimates to unsampled locations through the model [1]. The trade-off is that validity depends entirely on whether the assumed model is correct, which may be difficult to verify. In practice, its use for map accuracy assessment has remained limited [2]. National Forest Inventories (NFIs) represent a valuable source of high-quality, field-verified reference data. In cases where the inventory sampling design aligns with the target population and class definitions of the mapped product, for example, national land cover assessments using compatible forest definitions, NFI data can support probability-sampling-based validation. However, in practice, NFIs are also used to assess products with different thematic definitions, spatial extents, or target classes, such as global land cover products. In these situations, the original sampling design may no longer correspond to the validation objective, and the resulting assessment should be interpreted with appropriate caution.

In addition, NFI data are typically not openly accessible and often require formal collaboration for use in validation, particularly when applied to global products. Nevertheless, when appropriately matched to the validation objective and sampling framework, NFI-based assessments provide valuable and independent evidence of map quality and remain among the most reliable sources of reference data currently available.

For biomass mapping specifically, the CEOS WGCV Land Product Validation subgroup has developed dedicated protocols that address the particular challenges of validating continuous-field products, including the use of plot-level field measurements, allometric uncertainty, and spatial scaling from plot to pixel [3]. These protocols recommend combining field inventory data with airborne LiDAR transects to bridge the scale gap between ground plots and satellite-derived estimates. Even under such protocols, however, spatial coverage remains fundamentally limited, and the resulting accuracy characterization applies only to the sampled regions. More broadly, when probability sampling is not feasible, validation efforts should still adhere to general principles of transparency and rigor: reference data sources and their limitations should be clearly documented, the distinction between verification and validation should be made explicit, and results should be interpreted as indicative rather than as unbiased estimates of map accuracy.

The preceding sections have traced the full arc of the EO mapping pipeline, from the processing of satellite observations through to a validated, disseminated product. At each stage, we have highlighted choices that are easy to get wrong and hard to diagnose after the fact — from the silent differences between data providers, through the spatial structure of training splits, to the gap between algorithm-derived uncertainty and genuine accuracy assessment. We conclude by summarising the key recommendations that emerge from this analysis, examining the structural and institutional challenges that the community must confront as ML-based mapping matures from a research activity into an operational capability, and acknowledging the limitations of the present work.

[1]
S. V. Stehman and G. M. Foody, Key issues in rigorous accuracy assessment of land cover products,” Remote Sensing of Environment, vol. 231, p. 111199, 2019, doi: 10.1016/j.rse.2019.05.018.
[2]
A. Tyukavina, S. V. Stehman, A. H. Pickens, P. Potapov, and M. C. Hansen, Practical global sampling methods for estimating area and map accuracy of land cover and change,” Remote Sensing of Environment, vol. 324, p. 114714, 2025.
[3]
L. Duncanson et al., Global Aboveground Biomass Product Validation Best Practices Protocol,” in Best practice protocol for satellite derived land product validation, L. Duncanson, M. Disney, J. Armston, D. Minor, F. Camacho, and J. Nickeson, Eds., Land Product Validation Subgroup (WGCV/CEOS), 2020, p. 222. doi: 10.5067/doc/ceoswgcv/lpv/agb.001.