Design-based accuracy assessment

A statistically rigorous map accuracy assessment in a design-based inference framework relies on three essential components [1]: - the sampling design used to select the reference sample, - the response design used to obtain the reference for each sampling unit, - the estimation and analysis procedures.

These components must be considered jointly [2], [3], [4]. Across all stages, a clear and consistent problem definition is critical and should precede the validation process. In particular, the definitions used in sampling and labeling (such as what constitutes “forest” or the boundaries between land-cover classes) must be explicitly defined and aligned with the broader conceptual framing of the map product. ## Sampling design Sampling design determines how locations are selected for validation across the mapped domain. Design-based inference derives its validity from the randomness of the sample selection process rather than from model assumptions [2]. Under this framework, every location within the study area must have a known, non-zero probability of being selected. This requirement is known as the inclusion probability criterion [4], [5]. A wide range of probability sampling strategies can be applied, including simple random, systematic, and stratified approaches. The choice of sampling design, as well as the definition of strata, directly influences the statistical properties of the resulting accuracy estimates [6]. For example, different designs may affect the balance between the precision of user’s accuracy and producer’s accuracy, particularly in the presence of class imbalance [7]. Among available options, stratified random sampling is widely recommended as an adequate general-purpose design [6], [8]. It allows better control over sample allocation across strata (e.g., land cover classes) while maintaining a probabilistic framework for unbiased inference.

For maps that include rare classes such as deforestation or urban expansion, buffer-based stratification methods can substantially improve estimation. [7] propose introducing an additional stratum around mapped change areas, increasing the likelihood of sampling near class boundaries where omission errors concentrate. Sample size must balance statistical precision against the practical cost of reference labeling. Minimum sample sizes per class are generally recommended to ensure reliable per-class accuracy estimates, and formal procedures for determining sample size under stratified designs are well established [3], [6]. In practice, a widely adopted implementation is the stratified one-stage cluster approach [5], which combines stratification and clustering to balance statistical efficiency with the logistical constraints of human interpretation. The study area is first divided into strata based on attributes such as biomes, ensuring that rare but important classes (e.g., urban areas or water bodies) are sufficiently represented. Within each stratum, the area is partitioned into clusters or blocks, referred to as primary sampling units (PSUs). In a one-stage design, once a cluster is selected, every pixel within it is evaluated as part of the reference data [9]. Because clusters are selected from different strata at varying frequencies, each sample is assigned a weight based on its inclusion probability, enabling unbiased estimation of overall map accuracy and class-specific areas [10]. A concrete example of this approach is the multi-purpose Global Land Cover Validation dataset developed for the Copernicus Global Land Service [10], [11]. This dataset uses a global stratification that is independent of any land cover map and employs the Sentinel-2 UTM grid as its geographic base. It comprises more than 21 000 PSUs globally, each containing one hundred \(10 \times 10\) m reference pixels, and is updated annually to capture areas of change. The dataset has been used to validate multiple products, including the CGLS-LC100m and ESA WorldCover maps, demonstrating that CEOS WGCV Stage 4 validation, currently the highest validation stage achieved for satellite-derived land-cover products and characterized by sustained validation across product updates, is practically feasible at global scale. An alternative systematic approach uses a Discrete Global Grid System (DGGS) based on a hierarchical hexagonal tessellation of the globe, which minimizes geometric distortions inherent in regular latitude-longitude grids. A PSU is defined within each hexagon cell and reprojected into the local UTM zone to ensure alignment with the sensor pixels used for classification, guaranteeing equal-area sampling across the globe [4]. ## Response design Response design specifies how reference labels are assigned at the sampled locations. It encompasses the choice of spatial assessment unit (pixel, spatial window, or object-based), the source of reference information (field observations, high-resolution imagery, or crowdsourced data), the labeling protocol, and the definition of agreement between map and reference labels. Each of these choices must be explicitly documented [4].

In practice, reference labels are most commonly derived from visual interpretation of high-resolution imagery by trained experts. While widely regarded as the best available approach, this process is not fully reproducible: interpreter disagreement rates of 30

Several practices help mitigate these uncertainties. Reference data should be produced by interpreters who are independent from the map makers and unaware of map labels at sample locations, to avoid confirmation bias. Using multiple interpreters with consensus-based labeling further improves reliability [4]. Positional uncertainty can be addressed through the use of alternative labels derived from the spatial neighborhood of each sample unit, for instance by considering the majority class within a \(3\times3\) kernel around each reference pixel [9], [12]. Finally, reference data should be temporally matched to the map epoch [10]. ## Analysis Analysis refers to the statistical procedures used to derive accuracy metrics from the validation sample. For categorical maps, this typically involves constructing an error (confusion) matrix, which should be expressed in terms of area proportions rather than sample counts [4]. When sampling is stratified or otherwise non-uniform, unequal inclusion probabilities between strata must be accounted for, since sample sites are typically not allocated proportionally to strata areas [13], [14]. The inclusion probability for stratum \(h\) is defined as \(\pi_h = k_h / K_h\), where \(k_h\) is the number of sample sites in stratum \(h\) and \(K_h\) is the population size for that stratum. The estimation weight is then calculated as the inverse of the inclusion probability (\(1/\pi_h\)). These weights are used to construct an area-weighted confusion matrix from which unbiased estimates of overall and class-specific accuracies, as well as area proportions, can be derived [5], [11], [15]. More generally, when validation data are collected using a stratified sampling design, corresponding stratified estimators and appropriate sample weighting must be applied during analysis to obtain unbiased accuracy and area estimates [4]. User’s and producer’s accuracy (or equivalently, precision and recall), along with other metrics quantifying omission and commission errors, should be reported for individual classes in addition to overall accuracy. This is particularly important when a target class occupies a small proportion of the map, as high overall accuracy can mask poor performance on the class of interest [4]. In addition to traditional accuracy metrics, alternative decompositions of classification error, such as quantity disagreement and allocation disagreement [16], can provide complementary insight into the nature of map errors and are commonly adopted in the literature. For continuous-field products (e.g., tree cover fraction), adaptations of the standard confusion matrix framework, along with continuous accuracy metrics such as mean absolute error (MAE) and root mean square error (RMSE), are available [4], [5]. Accuracy metrics should be accompanied by measures of uncertainty, such as standard errors or confidence intervals, to reflect the variability inherent in the sampling process [17]. In practice, standard errors of accuracy metrics are often not reported, making it difficult for users to assess whether the provided estimates are reliable. Bootstrapping with replacement is commonly used to derive confidence intervals because it does not rely on parametric assumptions about the error distribution, although it does not correct systematic sampling bias in the validation design.

Reporting and transparency

Transparent reporting of validation procedures is essential for ensuring credibility, reproducibility, and appropriate interpretation of map products [17]. At a minimum, studies should document the sampling design (including strategy, sample size, and allocation), response design (reference data sources, labeling protocol, and uncertainty considerations), and analysis methods (accuracy metrics, area estimation procedures, and uncertainty measures such as confidence intervals). Any deviations from probability-based sampling or limitations in reference data should be clearly acknowledged. For products that are operationally updated, accuracy should be monitored for temporal stability, not only assessed at a single point in time. [10] propose to quantify stability by calculating an index from class accuracies at different moments in time, following the requirement that omission and commission error should each remain below 15

[1]
S. V. Stehman and R. L. Czaplewski, Design and analysis for thematic map accuracy assessment: fundamental principles,” Remote sensing of environment, vol. 64, no. 3, pp. 331–344, 1998.
[2]
S. V. Stehman, Statistical rigor and practical utility in thematic map accuracy assessment,” Photogrammetric Engineering and Remote Sensing, vol. 67, no. 6, pp. 727–734, 2001.
[3]
R. G. Congalton and K. Green, Assessing the Accuracy of Remotely Sensed Data: Principles and Practices, Third Edition. CRC Press, 2019. doi: 10.1201/9780429052729.
[4]
A. Tyukavina, S. V. Stehman, A. H. Pickens, P. Potapov, and M. C. Hansen, Practical global sampling methods for estimating area and map accuracy of land cover and change,” Remote Sensing of Environment, vol. 324, p. 114714, 2025.
[5]
B. Pengra, J. Long, D. Dahal, S. V. Stehman, and T. R. Loveland, A global reference database from very high resolution commercial satellite data and methodology for application to Landsat derived 30 m continuous field tree cover data,” Remote Sensing of Environment, vol. 165, pp. 234–248, Aug. 2015, doi: 10.1016/j.rse.2015.01.018.
[6]
P. Olofsson, G. M. Foody, M. Herold, S. V. Stehman, C. E. Woodcock, and M. A. Wulder, Good practices for estimating area and assessing accuracy of land change,” Remote Sensing of Environment, vol. 148, pp. 42–57, May 2014, doi: 10.1016/j.rse.2014.02.015.
[7]
P. Olofsson et al., Mitigating the effects of omission errors on area and area change estimates,” Remote Sensing of Environment, vol. 236, p. 111492, Jan. 2020, doi: 10.1016/j.rse.2019.111492.
[8]
S. V. Stehman, Sampling designs for accuracy assessment of land cover,” International Journal of Remote Sensing, vol. 30, no. 20, pp. 5243–5272, 2009, doi: 10.1080/01431160903131000.
[9]
P. Xu, N.-E. Tsendbazar, and M. Herold, Comparative validation of recent 10 m-resolution global land cover maps,” Remote Sensing of Environment, vol. 314, p. 114316, 2024, doi: 10.1016/j.rse.2024.114316.
[10]
N. Tsendbazar, M. Herold, L. Li, A. Tarko, and M. Duerauer, Towards operational validation of annual global land cover maps,” Remote Sensing of Environment, vol. 266, p. 112686, 2021, doi: 10.1016/j.rse.2021.112686.
[11]
N.-E. Tsendbazar et al., Developing and applying a multi-purpose land cover validation dataset for Africa,” Remote Sensing of Environment, vol. 219, pp. 298–309, Dec. 2018, doi: 10.1016/j.rse.2018.10.025.
[12]
N. E. Tsendbazar, S. de Bruin, B. Mora, L. Schouten, and M. Herold, Comparative assessment of thematic accuracy of GLC maps for specific applications using existing reference data,” International Journal of Applied Earth Observation and Geoinformation, vol. 44, pp. 124–135, Feb. 2016, doi: 10.1016/j.jag.2015.08.009.
[13]
P. Olofsson et al., A global land-cover validation data set, part I: fundamental design principles,” International Journal of Remote Sensing, vol. 33, no. 18, pp. 5768–5788, Mar. 2012, doi: 10.1080/01431161.2012.674230.
[14]
J. D. Wickham, S. V. Stehman, J. A. Fry, J. H. Smith, and C. G. Homer, Thematic accuracy of the NLCD 2001 land cover for the conterminous United States,” Remote Sensing of Environment, vol. 114, no. 6, pp. 1286–1296, 2010, doi: 10.1016/j.rse.2010.01.018.
[15]
P. Olofsson, G. M. Foody, S. V. Stehman, and C. E. Woodcock, Making better use of accuracy data in land change studies: Estimating accuracy and area and quantifying uncertainty using stratified estimation,” Remote Sensing of Environment, vol. 129, pp. 122–131, Feb. 2013, doi: 10.1016/j.rse.2012.10.031.
[16]
R. G. Pontius and M. Millones, Death to Kappa: birth of quantity disagreement and allocation disagreement for accuracy assessment,” International Journal of Remote Sensing, vol. 32, no. 15, pp. 4407–4429, Aug. 2011, doi: 10.1080/01431161.2011.552923.
[17]
S. V. Stehman and G. M. Foody, Key issues in rigorous accuracy assessment of land cover products,” Remote Sensing of Environment, vol. 231, p. 111199, 2019, doi: 10.1016/j.rse.2019.05.018.