Learning & representations · Research methods
Compare HybridSN, SpectralFormer and MambaHSI fairly: input context, whole-image access, pretraining, tuning, seeds and measured compute. Includes a protocol builder.

A CNN, a Transformer and a Mamba model can all classify hyperspectral pixels, but their reported scores may describe different experiments. One may receive a small neighbourhood, another a spectral sequence, and another an entire scene. The training labels, preprocessing, pretraining corpus, search effort and inference workload can differ as well. An architecture name does not settle those differences.
This article compares three concrete reference designs: HybridSN, SpectralFormer and MambaHSI. The goal is to design an interpretable experiment, rather than combine numbers from unrelated result tables into a leaderboard. There are no newly measured head-to-head accuracy results here. The interactive tool computes only input geometry, logical data volume and explicitly limited tensor-allocation arithmetic.
A useful comparison statement has a scope: which implementations performed best on which scenes, with which information and resource budget? Keep that sentence narrower than “Mamba beats Transformers” or “CNNs are obsolete.” The protocol builder makes the missing parts visible and exports them as an unfinished, reviewable experiment plan.
HybridSN combines a spectral–spatial 3D convolutional stage with a 2D spatial stage. It is a concrete hybrid CNN baseline, not a stand-in for every convolutional architecture. The paper version used here is arXiv:1902.06701v3, revised 3 July 2019. Its paper-linked implementation provides a 25 × 25 patch example and applies PCA to the input cube. The neighbourhood, retained components and fitting scope belong in the experimental specification. [1] [4]
SpectralFormer introduces group-wise spectral embeddings and cross-layer adaptive fusion, with both pixel-wise and patch-wise inputs. In the official code, spectral bands supply the token positions; a spatial patch contributes values within each token’s embedding input. The number of spatial pixels in the patch is therefore not automatically the attention sequence length. The paper version is arXiv:2107.02988v2, revised 20 November 2021; the journal citation is TGRS 2022. [2] [5]
MambaHSI combines spatial and spectral Mamba blocks with adaptive fusion. Its spatial branch is designed to model pixel-level interactions across a whole image, while the spectral branch processes groups of embedded features. That whole-image access is central to its comparison with patch-based systems. The journal article is from 2024; arXiv:2501.04944v1 was submitted on 9 January 2025. These dates identify the same work, rather than two model generations. [3]
A matched-information experiment gives every candidate the same eligible source data, labels, spectral representation and spatial input extent. Match the train/validation/test manifest, permitted pretraining data, augmentation information, selection metric and total search budget. Keep model-appropriate hyperparameters available within that budget. Equal learning rates or identical optimiser settings can unfairly handicap a model; equal opportunities are more meaningful than identical values.
This control has a cost. Restricting a whole-image method to isolated patches changes the method. Record the adapter, centre-pixel readout, padding, scan reset and any loss changes; describe it as an adapted baseline. If a native implementation cannot express the agreed input contract, say so. Do not quietly label the modified experiment an exact reproduction of the published system.
A native-pipeline experiment preserves each method’s intended context and preprocessing, while disclosing all differences. It answers an operational question about configured systems. It cannot cleanly attribute an advantage to a backbone when one system also sees more of the scene or receives extra pretraining. Reporting both comparisons makes those trade-offs useful: the matched track controls information, while the native track tests practical configurations.
Interactive research checklist · no training
Change the information regime, inspect input footprints, and export a protocol you can finish before running a real experiment.
These are explicitly chosen teaching assumptions. No head-to-head accuracy, latency or peak-memory results have been measured.
Native-inspired context sizes are 25 × 25, 7 × 7 and one full scene. The common patch control applies only in matched mode. Every format uses the same selected retained bands in this accounting; real native preprocessing needs a separate declaration.
| Illustrative format | One call | Raw input buffer | Naive whole-map input appearances |
|---|---|---|---|
| CNN / HybridSN-inspired25 × 25 patch | 16 patches16 prediction targets | 1.83 MiB30,000 values per patch | 491,520,000scalar input appearances |
| SpectralFormer-inspired25 × 25 patch | 16 patches16 prediction targets | 1.83 MiB30,000 values per patch | 491,520,000scalar input appearances |
| MambaHSI-inspired25 × 25 patch | 16 patches16 prediction targets | 1.83 MiB30,000 values per patch | 491,520,000scalar input appearances |
Naive whole-map accounting assumes one same-size zero-padded patch per scene pixel. It counts logical input appearances, including padding, rather than disk reads, MACs or simultaneously allocated memory. Efficient implementations can reuse input and features. Whole-image calls and patch calls produce different numbers of targets.
L = 49 (bands + one class token). Per-sequence component arithmetic, independent of the batch field.
These are different tensor components, not total model memory or a speed ranking. The state buffer is not a Mamba training-memory estimate. FlashAttention need not materialize the full score matrix. Actual activations, Q/K/V, convolutions, scan buffers, gradients, optimiser states, kernel workspace and batches are omitted.
Proposed upper bound: 39 single-device runs, 39.0 device-hours across three models. Each model gets 8 one-seed search trials plus 5 final-training seeds, at most 60 minutes per run. This is an unexecuted budget, not measured compute or a cost quote.
The exported JSON leaves dataset identity, split hashes, label budget, model revisions and hardware empty for you to supply. The example seeds are 100–104. No experiment is submitted or started.
Static default example. Enable JavaScript to change assumptions and export.
{
"schemaVersion": "hsi-comparison-protocol-1.0",
"status": "unexecuted-protocol-and-exact-toy-accounting",
"notABenchmark": true,
"settings": {
"mode": "matched",
"inspect": "cnn",
"split": "scene",
"target": "hidden",
"pretrain": "none",
"scaling": "train",
"side": 128,
"bands": 48,
"patch": 25,
"batch": 16,
"dtype": "fp32",
"trials": 8,
"seeds": 5,
"minutes": 60,
"axis": "spectral",
"heads": 4,
"dim": 64,
"state": 16
},
"question": "Compare adapted implementations under a common input extent and data regime.",
"data": {
"datasetId": null,
"sceneSplitManifestSha256": null,
"splitUnit": "scene",
"targetCovariatesAtTraining": "hidden",
"preprocessingFitScope": "train",
"pretraining": "none",
"inputBands": 48,
"labelBudgetPerClass": null,
"trainValidationTestIds": null,
"testLabelsUsedForSelection": false
},
"models": [
{
"id": "cnn",
"label": "CNN / HybridSN-inspired",
"context": "25 × 25 patch",
"sourceRevision": null,
"adaptationRequired": true,
"inputAccounting": {
"id": "cnn",
"name": "CNN / HybridSN-inspired",
"dense": false,
"context": "25 × 25 patch",
"patch": 25,
"inputValuesPerSample": 30000,
"inputBytesPerCall": 1920000,
"predictionTargetsPerCall": 16,
"callSamples": 16,
"naiveDenseMapInputValues": 491520000,
"overlap": {
"first": 625,
"second": 625,
"intersection": 425,
"union": 825,
"fraction": 0.68
}
}
},
{
"id": "transformer",
"label": "SpectralFormer-inspired",
"context": "25 × 25 patch",
"sourceRevision": null,
"adaptationRequired": true,
"inputAccounting": {
"id": "transformer",
"name": "SpectralFormer-inspired",
"dense": false,
"context": "25 × 25 patch",
"patch": 25,
"inputValuesPerSample": 30000,
"inputBytesPerCall": 1920000,
"predictionTargetsPerCall": 16,
"callSamples": 16,
"naiveDenseMapInputValues": 491520000,
"overlap": {
"first": 625,
"second": 625,
"intersection": 425,
"union": 825,
"fraction": 0.68
}
}
},
{
"id": "mamba",
"label": "MambaHSI-inspired",
"context": "25 × 25 patch",
"sourceRevision": null,
"adaptationRequired": true,
"inputAccounting": {
"id": "mamba",
"name": "MambaHSI-inspired",
"dense": false,
"context": "25 × 25 patch",
"patch": 25,
"inputValuesPerSample": 30000,
"inputBytesPerCall": 1920000,
"predictionTargetsPerCall": 16,
"callSamples": 16,
"naiveDenseMapInputValues": 491520000,
"overlap": {
"first": 625,
"second": 625,
"intersection": 425,
"union": 825,
"fraction": 0.68
}
}
}
],
"selection": {
"metric": "validation macro-F1",
"checkpointRule": "best validation metric within fixed cap; first checkpoint on ties",
"testSet": "evaluate once after configuration and checkpoint selection",
"search": {
"trialCountPerModel": 8,
"searchSeedsPerTrial": 1,
"finalSeeds": 5,
"perRunCapMinutes": 60,
"upperBoundRunsThreeModels": 39,
"upperBoundDeviceHoursThreeModels": 39
},
"seedList": [
100,
101,
102,
103,
104
],
"seedNote": "Example final-training seeds only. Publish separate split/search/initialisation seeds and all runs."
},
"compute": {
"hardware": null,
"softwareLock": null,
"precision": "fp32",
"trainingCapMinutes": 60,
"latencyProtocol": {
"warmup": 20,
"measuredRepeats": 100,
"report": [
"median",
"p95",
"IQR"
],
"scope": [
"model-only",
"full-scene end-to-end"
],
"include": [
"preprocessing",
"patch extraction or tiling",
"host-device transfers",
"stitching"
],
"acceleratorSynchronization": true,
"coldCompileTimeSeparate": true
},
"genericAllocationIllustration": {
"tokens": 49,
"axis": "spectral",
"bytesPerScalar": 4,
"materializedScoresBytes": 38416,
"oneTokenFeatureBytes": 12544,
"oneStreamingStateBytes": 4096
}
},
"outputsRequired": [
"per-class precision/recall/F1",
"macro-F1",
"overall accuracy",
"confusion matrix",
"per-scene results",
"paired seed results",
"training/search time",
"whole-scene latency",
"peak allocated and reserved device memory",
"host memory",
"failed and OOM runs"
],
"limitations": [
"Matched input extent requires adapted models. Patch-restricting MambaHSI changes its whole-image method; this is not an exact paper reproduction.",
"Scene-disjoint is a design intention here. Before running, verify scene IDs, provenance, duplicates and pretraining contamination."
],
"measurements": {
"accuracy": null,
"latency": null,
"peakMemory": null
},
"unfilledRequiredFields": [
"datasetId",
"sceneSplitManifestSha256",
"labelBudgetPerClass",
"trainValidationTestIds",
"all model sourceRevision values",
"hardware",
"softwareLock"
]
}For a strictly inductive, unseen-scene question, reserve evaluation scenes before fitting anything. Training may use complete training scenes; inference may use a complete held-out scene if that is how deployment works. Merely using a whole image at inference does not imply training leakage. What matters is when each scene becomes visible and which computations are allowed to depend on it.
For a transductive within-scene question, training can explicitly access unlabeled evaluation pixels. That can be a legitimate setting, but it is a different claim. Masking their labels in the loss does not remove their covariates from a whole-image forward pass, neighbourhood features, normalisation statistics or self-supervised objective. Report both label access and input access instead of using “no test labels” as a complete description.
Preprocessing needs the same audit. The inspected HybridSN notebook fits PCA to the reshaped cube; the inspected SpectralFormer demo obtains per-band minima and maxima from the image. Those are source-code facts, not accusations about every use of these methods. For a new inductive comparison, move learned preprocessing into the permitted training scope and document the deviation from the demonstration scripts. [4] [5]
Two disjoint lists of labelled pixel centres can still read overlapping source pixels. Liang and colleagues study this issue for spectral–spatial evaluation and explain why random same-image sampling can give misleading comparisons. For a claim about new regions or scenes, use an appropriate spatial or scene-level split and inspect the full processing footprint, including preprocessing filters. [7]
The lab’s toy geometry makes one part of that issue exact. Two 25 × 25 windows, centred eight columns apart at interior locations, share 17 × 25 = 425 of their 625 source pixels: 68%. Two 7 × 7 windows at those centres share none. Two full-image inputs share the entire 64 × 64 diagnostic grid. These are source-access counts, not model accuracy, effective receptive fields or proof that labels leaked.
Nonoverlapping footprints are helpful but do not establish statistical independence. Nearby fields, repeated specimens and overlapping acquisitions may remain related. A spatial buffer should account for every operation that mixes neighbours, and a scene holdout should match the intended geography, sensor, date or specimen transfer. Keep the richer split discussion separate from architecture claims, and publish the exact selected IDs.
Changing the retained bands, band order, normalisation, PCA or wavelength resampling changes the learning problem. A PCA component index is not a wavelength. If a baseline receives projected components while another receives ordered bands, record that difference rather than describing both inputs only as “30 channels.” In a controlled track, begin with a shared, training-fitted representation that each adapted model can accept, then add preprocessing ablations.
Pretraining can add both information and computation. A from-scratch model and a pretrained model may be a reasonable deployment comparison, but their difference does not isolate architecture. Record corpus identity, scene overlap checks, available wavelengths, objective, checkpoint, frozen versus fine-tuned layers and adaptation budget. Shared corpus access still does not make objectives or realised training cost identical.
Keep a from-scratch or common-corpus track when the scientific claim needs it. A pretrained checkpoint whose data provenance cannot be audited should carry that uncertainty into the conclusion. The existing foundation-model guide covers broader transfer questions; here, the practical requirement is to account for these resources in the comparison, including any evaluation data used before supervised training.
Choose the selection rule before viewing test results. For example, select a configuration and checkpoint using validation macro-F1, with a fixed tie rule and training cap, then evaluate the locked choice on the test set. Do not select the best test epoch, best test seed or most favourable scene after the fact. Use the same class definitions and label budgets across candidates.
The builder proposes eight one-seed search trials and five final-training seeds per model. With a one-hour cap per run, three models have an upper budget of 39 single-device hours. These are editable planning choices, not a universal recommendation, an executed workload or a cloud price. Equal trial counts and equal time answer different resource questions; record both, along with unsuccessful and out-of-memory trials.
Split seeds and training seeds measure different variation. Reuse paired split manifests across methods, list all training runs, and report per-scene results plus paired differences. Five seeds is only an example; it cannot manufacture five independent scenes. PyTorch also warns that identical seeds do not guarantee identical results across releases, devices or platforms. Pin the software environment and record deterministic-kernel settings and any performance cost. [9]
“Linear” and “quadratic” describe scaling under stated assumptions, not measured runtime for every input. Selective state-space sequence processing in Mamba motivates linear scaling with sequence length; it does not establish a universal HSI accuracy or latency advantage. Width, state size, number of passes, projections and implementation constants still matter. [8]
First identify the tokens. For the inspected SpectralFormer implementation, attention runs over band-derived tokens plus a class token. In MambaHSI’s spatial branch, sequence length follows the number of spatial pixels at that stage. Comparing one formula with L equal to bands against another with L equal to scene pixels is not a same-workload comparison. Likewise, spatial tiling, spectral grouping and patch extraction change the actual work. [3] [5] [6]
The resource panel intentionally computes only three named components: hL²s bytes for one materialized attention-score tensor, Lds bytes for one feature tensor, and dns bytes for one hypothetical streaming-state buffer. They are not interchangeable totals. FlashAttention uses an IO-aware exact-attention implementation that avoids the naive full-score storage pattern, so the first figure is not an unavoidable memory requirement of every Transformer. [10]
A useful latency report defines a workload: one isolated centre prediction, a batch of patches, or a complete output map for a specified scene. A whole-image call can produce thousands of predictions while a patch call produces a few. Comparing milliseconds per call without the output count is therefore misleading. Report both model-only inference and a full-scene end-to-end path.
For that end-to-end path, include preprocessing, patch extraction or tiling, data transfers and stitching. Disclose whether data are cached, whether compilation is already complete, and how boundaries are handled. Use the same device, precision policy, software stack where feasible, and stated CPU thread count. Synchronise asynchronous accelerator work before measuring completion, perform warmups, and report repeated timings with median and dispersion. PyTorch’s benchmark utilities explicitly handle these measurement issues. [11]
Peak device memory also needs a definition: allocated versus reserved memory, training versus inference, batch size, input shape and whether optimiser state already exists. Host memory and preprocessing buffers may be the actual bottleneck. MAC or FLOP counts require a convention and coverage of custom operators. Unsupported scan operators must not silently count as free. The lab has no timing engine and does not claim to estimate these quantities.
With its defaults, a 25 × 25 × 48 FP32 patch contains 30,000 scalar values and needs 120,000 bytes. Sixteen patches need 1,920,000 bytes, about 1.83 MiB, for that raw input buffer alone. In native-inspired mode, a 128 × 128 × 48 FP32 full-scene input needs 3,145,728 bytes, exactly 3 MiB, and offers 16,384 prediction targets. Neither number includes model activations or weights.
For a naive dense map assembled from one same-size patch per output pixel, the logical input volume is H × W × p² × B values. The default is 491,520,000 scalar appearances. This counts repeated values and zero-padding positions. It is not a statement about actual disk reads, simultaneous RAM allocation or convolutional work; streaming and feature reuse can change those costs.
The generic spectral example uses L = 48 + 1 = 49, four heads and four bytes per scalar. Its single score tensor has 38,416 bytes. Doubling L would quadruple that component; doubling precision bytes doubles it. The downloaded Python script independently reproduces these counts and enumerates the overlap geometry on a CPU, without external dependencies, model downloads or GPU jobs. Its purpose is to make assumptions checkable.
A completed experiment should show per-class precision, recall and F1, macro-F1, overall accuracy, confusion matrices and scene-level behaviour. Where fine structures matter, add a defined boundary or small-object assessment. A high aggregate score can hide the classes or regions that motivated the application. Link each result to its configuration, split hash and checkpoint.
Report the chosen model’s useful trade-offs: accuracy under the stated information regime, training and search cost, memory, full-scene latency, deployment constraints and sensitivity to context. Include failed runs and adaptations. If results differ between matched and native tracks, discuss which information or implementation choices may explain the difference; a targeted ablation is more informative than a family-wide verdict.
This article stops before that empirical conclusion. No HybridSN, SpectralFormer or MambaHSI training was run for this deliverable, and no head-to-head accuracy, runtime or peak-memory measurements are presented. The exported protocol remains marked unexecuted, its measurement fields stay empty, and required dataset, revision, label-budget and hardware fields must be completed before it can describe a real study.
No universal winner is established here. Results depend on the concrete implementation, data, information access, tuning and deployment workload; this article does not run a head-to-head benchmark.
No. A model trained on separate scenes can use the full held-out scene at inference. Evaluation covariates seen during training or preprocessing fitting must be separately disclosed.
Only after establishing compatible datasets, labels, splits, preprocessing, context and evaluation procedures. The article avoids cross-paper numerical ranking.
Restricting the whole-image design to isolated patches changes it. Document the adaptation and report it separately from a native-pipeline reproduction.
No. It counts raw inputs and explicitly named tensor components, omitting actual model state, activations, kernel workspace and runtime. Its streaming-state illustration is not a Mamba training-memory estimate.
No. It downloads an unexecuted plan with null measurement and provenance fields. The CPU script performs only arithmetic and toy geometry; it does not train or download models.