Synthetic comparison of one missing rare patch and twelve false rare islands, each with 97.92% overall accuracy but different rare-class precision and recall.
Original, deterministic synthetic figure. Both maps have twelve incorrect labels; their rare-class errors differ. No measured HSI data or benchmark results are shown.

Choose the failure you need to detect

Two classification maps can receive the same overall accuracy and fail in different ways. One can omit a rare material; another can detect all of it while inventing scattered occurrences elsewhere. A third can preserve class areas while putting their boundaries in the wrong place. Each failure changes what the map is useful for.

This guide concerns hard, single-label hyperspectral classification maps. It does not evaluate reconstructed spectra, calibrated reflectance, abundance estimates or probability calibration. Its aim is to make a map-quality report explain what was retained, missed, added and displaced.

The smoothing example explains one mechanism that removes small regions. Here, a separate synthetic workbench holds the reference fixed and introduces controlled errors. No spectral measurements, trained model or real benchmark result is represented.

A metric must retain its evaluated support, averaging and assumptions.Reference + evaluation mask to Count class disagreements: compare; Predicted labels to Count class disagreements: compare; Count class disagreements to Per-class + macro scores: summarise; Predicted labels to Boundary + region checks: locate errors; Reference + evaluation mask to Boundary + region checks: reference geometry; Per-class + macro scores to Report support + assumptions: with denominators; Boundary + region checks to Report support + assumptions: with conventionsConceptual relationshipsReference +evaluation maskPredicted labelsCount classdisagreementsPer-class + macroscoresBoundary + regionchecksReport support +assumptionscomparecomparesummariselocate errorsreference geometrywith denominatorswith conventions
Conceptual illustration. A metric must retain its evaluated support, averaging and assumptions.

Read the confusion matrix before the headline score

Use an explicit orientation: reference classes are rows and predicted classes are columns. Entry Cᵢⱼ counts evaluated pixels whose reference is i and prediction is j. Row totals give reference support; column totals give predicted support. Overall accuracy (OA) is the diagonal sum divided by the total evaluated support. A transposed matrix is valid too, but changes how precision and recall are read. [1]

For a selected class, true positives (TP) lie on its diagonal. False negatives (FN) are the other entries in its reference row; false positives (FP) are the other entries in its prediction column. Recall is TP/(TP+FN), precision is TP/(TP+FP), F1 is 2TP/(2TP+FP+FN), and class IoU is TP/(TP+FP+FN). Each answers a different question. [2] [6]

In remote-sensing terminology, producer’s accuracy corresponds to reference-class recall, while user’s accuracy corresponds to mapped-class precision when the same properly weighted matrix is used. State the orientation rather than assuming a reader will recognise the convention. [4]

Exactly calculated synthetic presets on the same 576-pixel reference. Nothing here is a trained-model benchmark.
Controlled errorIncorrect pixelsOARare precisionRare recall
Miss one rare patch1297.92%100.00%50.00%
Add 12 false rare islands1297.92%66.67%100.00%
Shift Region by one pixel2495.83%100.00%100.00%

Hold the scene fixed and change the error

The workbench contains 576 equal-area pixels: 432 Background, 120 Region and 24 Rare. Rare occupies two separate 3 × 4 patches, together 4.17% of the scene. These are arbitrary class names in a complete synthetic reference. They carry no claim about actual land cover.

The default removes one rare patch, causing 12 false negatives. OA is 564/576, or 97.92%, while rare-class recall is 12/24, or 50%. Choose “False islands” instead: 12 background pixels become rare-class predictions. OA remains 97.92%; rare recall rises to 100%, but precision falls to 24/36, or 66.67%. Equal error counts have concealed opposite operational problems.

Choose “Shift boundary” to translate only Region one pixel to the right. Its mapped area stays at 120 pixels, but 12 reference Region pixels are missed and 12 Background pixels are added to Region. OA is 95.83% and Region IoU is 108/132, or 81.82%. Correct mapped area alone has not established correct location.

Synthetic experiment · 576 equal-area pixels

One map, three error mechanisms

Compare a missed rare patch, false rare-class islands and a shifted Region boundary. Every result below is calculated from the displayed label maps.

BackgroundRegion (hatching)Rare (dots)Incorrect label
Hyperspectral Imaging & AI Blog | Muhammad HusnainSynthetic 24 by 24 map. The tables below give the complete class counts; crosses mark disagreements.
Reference432 Background · 120 Region · 24 Rare
Hyperspectral Imaging & AI Blog | Muhammad HusnainSynthetic 24 by 24 map. The tables below give the complete class counts; crosses mark disagreements.
PredictionControlled errors only; no model
Hyperspectral Imaging & AI Blog | Muhammad HusnainSynthetic 24 by 24 map. The tables below give the complete class counts; crosses mark disagreements.
Disagreements12 incorrect pixels

564 / 576 correct pixels. Overall accuracy 97.92%. Rare-class precision 100.00%; recall 50.00%.

Overall accuracy97.92%
Macro recall (balanced)83.33%
Macro F188.43%
Mean class IoU82.43%
Pixel counts. Rows: reference. Columns: prediction.
Reference ↓ / Predicted →BackgroundRegionRareSupport
Background43200432
Region01200120
Rare1201224
Per-class scores. All three classes are included.
ClassReference nPredicted nPrecisionRecallF1IoU
Background43244497.30%100.00%98.63%97.30%
Region120120100.00%100.00%100.00%100.00%
Rare2412100.00%50.00%66.67%50.00%

Macro precision: 99.10% across 3 defined classes. Macro recall, F1 and IoU each include all three reference classes in this scene.

Boundary and connected-region diagnostics

Region mask IoU: 100.00%. Region Boundary IoU: 100.00% (40 / 40 band pixels; width 1 pixel).

Rare-class components (4-neighbour): reference 2, prediction 1. Component sizes in pixels: reference [12, 12]; prediction [12].

The inner band is mask minus square-neighbourhood erosion, with outside-image background. These connected components are raster regions, not labelled physical objects.

What if the rare class had a different prevalence?

At a hypothetical 5% rare-class prevalence, overall accuracy would be 97.50% and rare-class precision 100.00%. Macro recall remains 83.33%. These are reweighted proportions, not additional evaluated pixels.

Reference-row conditional error rates stay fixed. The non-rare remainder is divided Background:Region = 18:5. The displayed maps and count matrix stay unchanged. Try the False islands preset to see how prevalence changes precision. This is neither a field-area estimator nor a domain-shift guarantee.

Calculation conventions and limits

All 576 pixels are labelled and evaluated; no missing-data mask is used. Three fixed classes are included. Undefined ratios are displayed as Undefined and excluded from the relevant macro mean. A reference class with no predictions retains zero recall, F1 and IoU.

Errors are deterministic: remove rare pixels in row-major order, add isolated rare labels along the top and near-bottom rows, and translate Region to the right. No smoothing, spectral input, training, uncertainty interval or instance matching is performed.

Boundary bands use Chebyshev-distance square erosion of radius d. An empty band union is undefined. Region is the only class selected for the boundary diagnostic. Changing band width changes the measurement support.

Source definitions: precision, recall and averaging; Boundary IoU. The article explains sampling and reference uncertainty separately.

Separate equal-class summaries from area-weighted accuracy

Macro averaging gives each included class equal weight. In this article, balanced accuracy means the unadjusted mean of class recalls, matching scikit-learn’s definition. The missed-patch preset has recalls of 100%, 100% and 50%, so macro recall is 83.33%. The rare class matters as much as either larger class in this summary. [3]

Macro F1 averages the individual class F1 scores; it is not the F1 calculated from macro precision and macro recall. Support-weighted averages weight classes by their reference counts. In a complete, single-label evaluation over all classes, support-weighted recall equals OA, so reporting both does not provide an independent diagnostic. [2]

A pixel-count OA is area-weighted for this complete equal-area grid. That interpretation does not automatically transfer to a benchmark’s selected labelled pixels. Oversampling rare classes changes the evaluation distribution. A class-balanced test set measures a different mixture from the mapped landscape, and equal numbers per class do not make it an area estimate.

For map-area inference, sampling and analysis must agree. Olofsson and colleagues recommend probability sampling, an appropriate reference-labelling protocol, area-proportion error matrices and uncertainty estimates. With map-class strata, a stratum’s mapped area fraction weights its estimated reference-class proportions. Naively treating all sampled labels as equally representative can misstate area accuracy. [4]

Change prevalence without pretending to collect more data

The prevalence control asks a counterfactual question: what would these conditional error rates imply if Rare occupied a different share of the population? It reweights each reference row, keeping the prediction distribution within that row fixed. Background and Region retain their relative 18:5 prevalence ratio within the non-rare remainder.

Select “False islands”, then lower hypothetical rare prevalence. The false-positive rate within Background stays fixed, yet false positives become a larger share of rare-class predictions. Rare precision falls. Macro recall stays fixed because every class’s conditional recall was held fixed; OA changes because the population mixture changed.

This is an arithmetic sensitivity analysis. It assumes that conditional errors remain stable under the prevalence change, which real sensor, geography or acquisition shifts may violate. The control neither changes the displayed 576-pixel map nor supplies an area estimate or confidence interval for an actual landscape.

Give boundary quality its own measurement

Class IoU compares all pixels of a class. A boundary error can therefore occupy a small fraction of a large region. Cheng and colleagues introduced Boundary IoU to compare near-boundary portions of predicted and reference masks, making boundary differences more visible in segmentation evaluation. Their paper studies object-centric segmentation; applying the idea to a class-level HSI mask requires stating the chosen protocol. [5]

Here, the inner band is the class mask minus its erosion by a square neighbourhood of radius d. Pixels beyond the image are treated as background. Boundary IoU is the intersection of reference and prediction bands divided by their union. The selected width is a fixed one, two or three pixels, not a tolerance selected to make a method look better.

For the one-pixel Region shift and a one-pixel band, the band intersection contains 18 pixels and the union contains 62. Boundary IoU is 29.03%, alongside mask IoU of 81.82%. The two measurements are deliberately sensitive to different supports. A wider band changes the question; retain the same width and border treatment when comparing maps.

Real evaluation also needs the spatial sampling distance and registration uncertainty. One pixel means different distances for different sensors. A strict band can penalise a small registration offset or an uncertain reference edge. Report sensitivity to justified widths, without tuning the final choice on the test result.

Treat connected regions as diagnostics, not object identities

The workbench reports connected components of Rare under four- or eight-neighbour connectivity. Four-neighbour connectivity joins edge-sharing pixels; eight-neighbour connectivity also joins diagonals. The generated false islands are separated even under eight-neighbour connectivity, so their counts are deliberately stable. Other maps need not be.

Component count and component-size distribution can expose fragmentation or scattered additions. They do not establish object precision or object recall. A connected land-cover patch can contain several physical objects, and one object can occupy several disconnected labelled patches.

An object-level assessment needs instance annotations and a matching rule: overlap or distance criteria, one-to-one assignment, and explicit treatment of splits and merges. Likewise, a total component count can agree while the locations are wrong. Use the count to guide inspection, alongside the map and class errors.

Make missing support and reference uncertainty visible

Precision is undefined when a class has no predicted pixels. Recall is undefined when it has no reference pixels. If a class appears in the reference but is never predicted, its recall, F1 and IoU are zero. When it appears in neither map, F1 and IoU are undefined too. [2]

Set missed rare pixels to 24 and false islands to zero. The table displays undefined rare precision and zero rare recall, F1 and IoU. This workbench excludes undefined values from each macro average and prints the precision denominator. That convention can make macro precision look reassuring when a class disappears, so the per-class table must remain visible. A library configured to replace undefined values with zero will give different averages.

Outside this synthetic scene, an unlabelled pixel is not automatically background. Define an evaluation mask, report its coverage and keep unknown regions out of both denominators unless they have an explicit reference label. Check spatial alignment, annotation date, mixed-pixel rules and class definitions before interpreting disagreement as model error.

Reference uncertainty and sampling uncertainty are separate. More sampled pixels cannot repair a systematically incorrect reference. For uncertainty intervals, describe the sampling or resampling unit and its dependence assumptions; neighbouring pixels should not simply be treated as independent replications. Record uncertain or disputed labels and any adjudication procedure. [4]

Build a report that supports a decision

Choose the target use before selecting the headline metric. Missing a scarce material, creating false detections and displacing a field boundary can have different consequences. A single weighted score is defensible only when its trade-offs are explicit; it should not hide the underlying errors.

For a fine-grained HSI classification report, keep the evaluation contract beside the results. Show reference and prediction maps at the same scale, the disagreement map, the confusion matrix and the class support. Add boundary or connected-region diagnostics when the task needs them, rather than treating every available metric as a new claim of quality.

  • Specify the class ontology, evaluation mask, spatial sampling distance, split design and available test support.
  • Report the matrix orientation, OA, per-class precision and recall, and at least one clearly defined macro summary.
  • Name class exclusions, averaging weights and zero-division conventions. Keep missed classes visible.
  • For spatial diagnostics, fix connectivity, boundary-band construction, width and image-edge handling.
  • Separate sample performance, area inference and uncertainty; explain which claim the design supports.
  • Inspect errors by scene, region or acquisition where the sample design permits. A pooled score can conceal a weak deployment setting.

Frequently asked questions

Is a high overall accuracy sufficient for rare-class mapping?

No. OA weights each evaluated pixel equally. Read the rare class’s support, precision and recall, and check whether evaluation sampling represents the intended population.

Is average accuracy the same as area-adjusted accuracy?

No. A mean of class recalls is class-balanced. Area-adjusted accuracy requires weights and analysis consistent with the map population and sampling design. Define any use of the abbreviation AA.

Does a zero prediction count mean 100% precision?

No. Precision has a zero denominator and is undefined. A reporting convention may replace it with a number, but that choice must be stated alongside recall and class support.

Can Boundary IoU replace per-class accuracy?

It measures agreement of specified boundary bands. It complements class-level errors and mask IoU; it does not measure every aspect of map quality.

Are connected components equivalent to objects?

No. Components depend on the raster labels and connectivity rule. Object-level scores require instance references and a matching protocol.

Are the workbench percentages research results?

They are exact calculations on a constructed 576-pixel label grid. They illustrate measurement behaviour, not the performance of an HSI method or dataset.

References and further reading

  1. scikit-learn: confusion_matrix (reference rows, prediction columns)
  2. scikit-learn: precision_recall_fscore_support (averages and undefined values)
  3. scikit-learn: balanced_accuracy_score
  4. Olofsson et al. (2014), Good practices for estimating area and assessing accuracy of land change
  5. Cheng et al. (CVPR 2021), Boundary IoU: Improving Object-Centric Image Segmentation Evaluation
  6. scikit-learn: Jaccard similarity coefficient / intersection over union

Reader feedback

— reads— comments

Reads since 1 October 2026: at least 15 seconds with the article visible, counted once per browser per day. Reactions are anonymous and can be changed.

Discuss this article

Ask a technical question, challenge an assumption or share evidence from your own work.

Your name and comment will be public after approval. New comments wait for approval. No email required. Keep it respectful and relevant; no personal information or spam. Plain text, up to 2,000 characters.

Owner: manage comments