
A cleaner map can be less faithful
A hyperspectral classification map assigns a material or land-cover label to each spatial sample. Readers often expect neighbouring pixels to form coherent regions. Removing isolated predictions can improve a noisy map, but visual tidiness does not establish correctness. A small genuine region can occupy exactly the shape that a smoothing operator treats as an error.
This article separates two ideas. Majority filtering changes categorical labels according to nearby votes. Representation oversmoothing in graph learning concerns features becoming less distinguishable after repeated propagation. They can both reduce useful distinctions, but they are different operations with different explanations. The public GCN analysis provides background for the second meaning [1].
We will use a deterministic toy map to inspect the first meaning. The reference labels are constructed for teaching, not taken from a field survey or produced by a trained hyperspectral model. That makes the example easy to reproduce and its limits easy to see.
Why a four-pixel region loses a local vote
Place a 2 × 2 minority region inside a uniform background. For each of its four pixels, a 3 × 3 neighbourhood contains four minority labels and five background labels. A simultaneous majority filter therefore replaces each minority label with the surrounding class. It does not inspect the spectrum or ask whether the small region is physically plausible.
The simultaneous update matters. Every output label must be computed from the same input map. Updating the input in place would let earlier changes affect later neighbourhoods and create an order-dependent algorithm. The code below uses a separate prediction list. At the outer image border, the window is clipped to valid pixels.
The filter has followed its specified rule correctly. The failure is a mismatch between that rule and the task. If genuine objects can be smaller than the window, local dominance is not sufficient evidence that an object is false. This does not prove every spatial model erases small objects; it demonstrates one precise failure mode.
| Measurement | What it asks | Toy result or limit |
|---|---|---|
| Overall accuracy | How many map labels are correct? | 98.96% |
| Tiny-class recall | How many true tiny-class pixels are retained? | 0% |
| Tiny-class precision | How many predicted tiny-class pixels are correct? | No positive predictions; convention must be stated |
| Connectivity | Does a relevant connection survive? | Not determined by accuracy alone |
| Visual smoothness | Does the map look coherent? | Not a correctness measurement |
Overall accuracy answers a different question
Our map has 24 × 16 = 384 pixels. If the four object pixels are the only errors, 380 pixels remain correct. Overall accuracy is therefore 380/384, or 98.96%. The tiny class has four reference pixels and zero correctly predicted pixels, so its recall is 0%. Both numbers describe the same output.
Accuracy counts all correct predictions equally in the global denominator. Recall for a class divides that class’s true positives by its true positives plus false negatives. Precision instead asks what fraction of predictions for the class are correct [2]. Reporting recall alone would not detect a method that labels many background pixels as tiny objects.
The example is deliberately extreme. It does not estimate the prevalence of real tiny objects in a dataset, nor establish a model’s performance. Its value is that it makes the denominator visible. A global summary can remain high while a scientifically important class disappears.
A RESEARCH QUESTION YOU CAN EXPLORE
Smoother isn’t always
more accurate.
A neat-looking map can hide a real object. Local smoothing favours whichever class dominates a neighbourhood, so small regions and narrow features can lose their identity.
These deterministic toy scenes apply a 3×3 majority filter to class labels. They illustrate a failure mode; they are not outputs of a trained model or a claim that every model behaves this way.
Four pixels. One genuine object.
The mint square is a real minority object. In a 3×3 neighbourhood, its four pixels lose the majority vote to the surrounding class.
WHY A GLOBAL SCORE CAN HIDE IT
A small denominator
changes the story.
Losing four pixels changes only 1% of this map, but loses 100% of the tiny object. A whole-map summary and an object-level measure answer different questions.
WHY HYPERSPECTRAL DETAIL MATTERS
Similar-looking
isn’t the same material.
Hyperspectral sensors measure many wavelength bands. Subtle spectral differences can separate related land-cover materials, while a spatial model or refiner may still favour the surrounding class.
Weak supervision, mixed pixels and uncertain boundaries make preserving those distinctions harder.
THE OTHER SIDE OF THE TRADE-OFF
Keep the signal.
Reject the false island.
A tiny region can also be noise. Preserving every small blob creates false positives; smoothing every blob destroys genuine objects. The aim is to use spectral evidence and context to distinguish the two.
The example motivates measuring fine classes and real spatial structures alongside global accuracy.
Explore object size, boundaries and fine classes
The interactive maps below let you inspect a tiny region, a narrow connection and a small fine-class patch. They use categorical majority filtering and expose the reference and transformed map together. White outlines identify the focus region so the failure is visible even after its label changes.
Try zero passes first, then increase the number of passes. Compare the number of retained focus pixels with the fraction of the entire map that remains unchanged. The latter is a change rate, not accuracy unless the starting map is a trusted reference. In this teaching example the starting labels are defined as the reference.
A thin connection shows why pixel counts do not capture every structural consequence. Losing a few pixels can separate two connected regions. Conversely, a method can preserve the total area of a class while relocating it. A task that depends on location or connectivity needs measurements that respond to those changes.
A small region can also be wrong
Preserving every isolated prediction is not a solution. Sensor noise, shadows, registration error and model uncertainty can create false islands. A map may contain both genuine small regions and incorrect small blobs. Their similar geometry prevents size alone from resolving the problem.
The measured spectrum provides additional evidence, but that evidence has limits. A spatial sample can contain more than one material, and spectra can change with illumination, acquisition conditions or material variability. The public unmixing overview describes why a mixed pixel should not automatically be treated as a pure material label [4].
Choose the evaluation target before choosing the clean-up operation. A region can be important because it represents a rare material, a disease spot, a narrow physical feature or a decision-relevant boundary. Those are application statements that need trustworthy reference information, not an argument for retaining every colourful dot.
Evaluate the effect without hiding the trade-off
Keep the classifier output and the transformed map. Compare them against the same held-out reference using overall and per-class metrics. Report how many pixels change, where the changes occur and whether small reference regions become detectable. Also inspect false small regions produced in the background.
Choose and document the connectivity rule, minimum region size and matching criterion before comparing object counts. Four-connected and eight-connected components can give different region counts. An object detection score also needs a clear location or overlap rule; it is not interchangeable with pixel recall.
Make the data split explicit. Spatially overlapping train and test patches can make performance optimistic even when label centres differ. The public sampling study explains why random pixel splits require special care for spectral–spatial methods [3]. Use an evaluation protocol that reflects the intended deployment question.
Practical best practices
Start with the unsmoothed output and inspect errors before introducing a transformation. Treat smoothing strength as a model-selection choice and tune it on validation data, rather than selecting it after examining the test labels.
Keep a per-class table and several representative views at the same scale. Compare the same geographic areas and use a stable palette. Record the operator, boundary handling, number of passes and any tie rule so another reader can reproduce the transformation.
For a real study, include uncertainty in the interpretation. A disputed boundary or partially mixed pixel may not support an unambiguous hard label. The useful question is whether the procedure improves the task under its stated evidence, not whether the map looks more polished.
What the example establishes
The toy calculation establishes that a 3 × 3 local majority rule can erase a 2 × 2 minority region while leaving overall accuracy close to 99%. It does not establish a new recovery method, a benchmark result or the behaviour of a particular neural architecture.
For a broader introduction to spectral measurements, continue with the HSI fundamentals guide. For supervision and evaluation terminology, read the weak-supervision guide.
Run the example
Prerequisite: Python 3. Examples use synthetic inputs to explain the calculation. Save the snippet as example.py and run python3 example.py.
from collections import Counter
width, height = 24, 16
reference = [0] * (width * height)
for y in (7, 8):
for x in (11, 12):
reference[y * width + x] = 1
prediction = []
for y in range(height):
for x in range(width):
votes = Counter(reference[yy * width + xx]
for yy in range(max(0, y-1), min(height, y+2))
for xx in range(max(0, x-1), min(width, x+2)))
prediction.append(votes.most_common(1)[0][0])
correct = sum(a == b for a, b in zip(reference, prediction))
retained = sum(a == b == 1 for a, b in zip(reference, prediction))
print(f'Overall accuracy: {100 * correct / len(reference):.2f}%')
print(f'Tiny-class recall: {100 * retained / sum(reference):.1f}%')
print(f'Object pixels retained: {retained}/4')
Verified output
Overall accuracy: 98.96% Tiny-class recall: 0.0% Object pixels retained: 0/4
Frequently asked questions
Does every HSI model oversmooth?
No. The example demonstrates a majority filter on labels, not every model or training objective.
Is a tiny label region always noise?
No. Its geometry alone cannot distinguish a genuine region from an incorrect prediction.
Why can accuracy stay high when an object disappears?
The object occupies a small part of the global pixel denominator.
Is graph oversmoothing the same as majority filtering?
No. One concerns feature propagation; the other is an explicit categorical label operation.
Can the toy example validate a recovery method?
No. Real-scene references, independent splits and false-positive measurements are also needed.
Are these results from unpublished research?
No. They are deterministic calculations on a constructed teaching map.
References and further reading
- Deeper Insights into Graph Convolutional Networks for Semi-Supervised Learning
- scikit-learn: classification metrics
- On the Sampling Strategy for Evaluation of Spectral-spatial Methods in Hyperspectral Image Classification
- Image Processing and Machine Learning for Hyperspectral Unmixing: An Overview and the HySUPP Python Package