Learning & representations · Research methods

CNNs, Transformers and Mamba for HSI: a fair comparison

Compare HybridSN, SpectralFormer and MambaHSI fairly: input context, whole-image access, pretraining, tuning, seeds and measured compute. Includes a protocol builder.

Original input-footprint diagram: two 25 by 25 windows overlap in 425 source pixels, two 7 by 7 windows do not overlap, and two whole-scene inputs share all 4096 pixels on a 64 by 64 toy grid.
Exact toy input geometry for centres eight columns apart. Window size is a chosen context example; this figure shows no learned influence, accuracy or architecture winner.

A fair comparison starts with information

A CNN, a Transformer and a Mamba model can all classify hyperspectral pixels, but their reported scores may describe different experiments. One may receive a small neighbourhood, another a spectral sequence, and another an entire scene. The training labels, preprocessing, pretraining corpus, search effort and inference workload can differ as well. An architecture name does not settle those differences.

This article compares three concrete reference designs: HybridSN, SpectralFormer and MambaHSI. The goal is to design an interpretable experiment, rather than combine numbers from unrelated result tables into a leaderboard. There are no newly measured head-to-head accuracy results here. The interactive tool computes only input geometry, logical data volume and explicitly limited tensor-allocation arithmetic.

A useful comparison statement has a scope: which implementations performed best on which scenes, with which information and resource budget? Keep that sentence narrower than “Mamba beats Transformers” or “CNNs are obsolete.” The protocol builder makes the missing parts visible and exports them as an unfinished, reviewable experiment plan.

Three reference designs, three ways to use context

HybridSN combines a spectral–spatial 3D convolutional stage with a 2D spatial stage. It is a concrete hybrid CNN baseline, not a stand-in for every convolutional architecture. The paper version used here is arXiv:1902.06701v3, revised 3 July 2019. Its paper-linked implementation provides a 25 × 25 patch example and applies PCA to the input cube. The neighbourhood, retained components and fitting scope belong in the experimental specification. [1] [4]

SpectralFormer introduces group-wise spectral embeddings and cross-layer adaptive fusion, with both pixel-wise and patch-wise inputs. In the official code, spectral bands supply the token positions; a spatial patch contributes values within each token’s embedding input. The number of spatial pixels in the patch is therefore not automatically the attention sequence length. The paper version is arXiv:2107.02988v2, revised 20 November 2021; the journal citation is TGRS 2022. [2] [5]

MambaHSI combines spatial and spectral Mamba blocks with adaptive fusion. Its spatial branch is designed to model pixel-level interactions across a whole image, while the spectral branch processes groups of embedded features. That whole-image access is central to its comparison with patch-based systems. The journal article is from 2024; arXiv:2501.04944v1 was submitted on 9 January 2025. These dates identify the same work, rather than two model generations. [3]

Run two comparisons if you need two answers

A matched-information experiment gives every candidate the same eligible source data, labels, spectral representation and spatial input extent. Match the train/validation/test manifest, permitted pretraining data, augmentation information, selection metric and total search budget. Keep model-appropriate hyperparameters available within that budget. Equal learning rates or identical optimiser settings can unfairly handicap a model; equal opportunities are more meaningful than identical values.

This control has a cost. Restricting a whole-image method to isolated patches changes the method. Record the adapter, centre-pixel readout, padding, scan reset and any loss changes; describe it as an adapted baseline. If a native implementation cannot express the agreed input contract, say so. Do not quietly label the modified experiment an exact reproduction of the published system.

A native-pipeline experiment preserves each method’s intended context and preprocessing, while disclosing all differences. It answers an operational question about configured systems. It cannot cleanly attribute an advantage to a backbone when one system also sees more of the scene or receives extra pretraining. Reporting both comparisons makes those trade-offs useful: the matched track controls information, while the native track tests practical configurations.

Interactive research checklist · no training

Build the comparison before choosing a winner

Change the information regime, inspect input footprints, and export a protocol you can finish before running a real experiment.

These are explicitly chosen teaching assumptions. No head-to-head accuracy, latency or peak-memory results have been measured.

64 × 64 toy source grid
First centreSecond centreShared input
25 × 25 patch: 425 of 625 source pixels shared (68.0%). Centres are eight columns apart on a separate 64 × 64 toy grid. This counts input access, not learned influence or label leakage.
MATCHED INPUT EXTENT · ADAPTED MODELS

What would this comparison establish?

  • Matched input extent requires adapted models. Patch-restricting MambaHSI changes its whole-image method; this is not an exact paper reproduction.
  • Scene-disjoint is a design intention here. Before running, verify scene IDs, provenance, duplicates and pretraining contamination.
Input, information access and experiment budget

Native-inspired context sizes are 25 × 25, 7 × 7 and one full scene. The common patch control applies only in matched mode. Every format uses the same selected retained bands in this accounting; real native preprocessing needs a separate declaration.

Input storage for the chosen scene dimensions and precision, excluding model memory
Illustrative formatOne callRaw input bufferNaive whole-map input appearances
CNN / HybridSN-inspired25 × 25 patch16 patches16 prediction targets1.83 MiB30,000 values per patch491,520,000scalar input appearances
SpectralFormer-inspired25 × 25 patch16 patches16 prediction targets1.83 MiB30,000 values per patch491,520,000scalar input appearances
MambaHSI-inspired25 × 25 patch16 patches16 prediction targets1.83 MiB30,000 values per patch491,520,000scalar input appearances

Naive whole-map accounting assumes one same-size zero-padded patch per scene pixel. It counts logical input appearances, including padding, rather than disk reads, MACs or simultaneously allocated memory. Efficient implementations can reuse input and features. Whole-image calls and patch calls produce different numbers of targets.

Inspect generic allocation components

L = 49 (bands + one class token). Per-sequence component arithmetic, independent of the batch field.

One materialized attention-score tensor37.52 KiB4 × 49² × 4 bytes
One token-feature tensor12.25 KiB49 × 64 × 4 bytes
One hypothetical streaming-state buffer4.00 KiB64 × 16 × 4 bytes

These are different tensor components, not total model memory or a speed ranking. The state buffer is not a Mamba training-memory estimate. FlashAttention need not materialize the full score matrix. Actual activations, Q/K/V, convolutions, scan buffers, gradients, optimiser states, kernel workspace and batches are omitted.

An experiment you can audit

Proposed upper bound: 39 single-device runs, 39.0 device-hours across three models. Each model gets 8 one-seed search trials plus 5 final-training seeds, at most 60 minutes per run. This is an unexecuted budget, not measured compute or a cost quote.

The exported JSON leaves dataset identity, split hashes, label budget, model revisions and hardware empty for you to supply. The example seeds are 100–104. No experiment is submitted or started.

Static default example. Enable JavaScript to change assumptions and export.

Download the CPU-only estimator · Download source and tests

Read the current protocol JSON
{
  "schemaVersion": "hsi-comparison-protocol-1.0",
  "status": "unexecuted-protocol-and-exact-toy-accounting",
  "notABenchmark": true,
  "settings": {
    "mode": "matched",
    "inspect": "cnn",
    "split": "scene",
    "target": "hidden",
    "pretrain": "none",
    "scaling": "train",
    "side": 128,
    "bands": 48,
    "patch": 25,
    "batch": 16,
    "dtype": "fp32",
    "trials": 8,
    "seeds": 5,
    "minutes": 60,
    "axis": "spectral",
    "heads": 4,
    "dim": 64,
    "state": 16
  },
  "question": "Compare adapted implementations under a common input extent and data regime.",
  "data": {
    "datasetId": null,
    "sceneSplitManifestSha256": null,
    "splitUnit": "scene",
    "targetCovariatesAtTraining": "hidden",
    "preprocessingFitScope": "train",
    "pretraining": "none",
    "inputBands": 48,
    "labelBudgetPerClass": null,
    "trainValidationTestIds": null,
    "testLabelsUsedForSelection": false
  },
  "models": [
    {
      "id": "cnn",
      "label": "CNN / HybridSN-inspired",
      "context": "25 × 25 patch",
      "sourceRevision": null,
      "adaptationRequired": true,
      "inputAccounting": {
        "id": "cnn",
        "name": "CNN / HybridSN-inspired",
        "dense": false,
        "context": "25 × 25 patch",
        "patch": 25,
        "inputValuesPerSample": 30000,
        "inputBytesPerCall": 1920000,
        "predictionTargetsPerCall": 16,
        "callSamples": 16,
        "naiveDenseMapInputValues": 491520000,
        "overlap": {
          "first": 625,
          "second": 625,
          "intersection": 425,
          "union": 825,
          "fraction": 0.68
        }
      }
    },
    {
      "id": "transformer",
      "label": "SpectralFormer-inspired",
      "context": "25 × 25 patch",
      "sourceRevision": null,
      "adaptationRequired": true,
      "inputAccounting": {
        "id": "transformer",
        "name": "SpectralFormer-inspired",
        "dense": false,
        "context": "25 × 25 patch",
        "patch": 25,
        "inputValuesPerSample": 30000,
        "inputBytesPerCall": 1920000,
        "predictionTargetsPerCall": 16,
        "callSamples": 16,
        "naiveDenseMapInputValues": 491520000,
        "overlap": {
          "first": 625,
          "second": 625,
          "intersection": 425,
          "union": 825,
          "fraction": 0.68
        }
      }
    },
    {
      "id": "mamba",
      "label": "MambaHSI-inspired",
      "context": "25 × 25 patch",
      "sourceRevision": null,
      "adaptationRequired": true,
      "inputAccounting": {
        "id": "mamba",
        "name": "MambaHSI-inspired",
        "dense": false,
        "context": "25 × 25 patch",
        "patch": 25,
        "inputValuesPerSample": 30000,
        "inputBytesPerCall": 1920000,
        "predictionTargetsPerCall": 16,
        "callSamples": 16,
        "naiveDenseMapInputValues": 491520000,
        "overlap": {
          "first": 625,
          "second": 625,
          "intersection": 425,
          "union": 825,
          "fraction": 0.68
        }
      }
    }
  ],
  "selection": {
    "metric": "validation macro-F1",
    "checkpointRule": "best validation metric within fixed cap; first checkpoint on ties",
    "testSet": "evaluate once after configuration and checkpoint selection",
    "search": {
      "trialCountPerModel": 8,
      "searchSeedsPerTrial": 1,
      "finalSeeds": 5,
      "perRunCapMinutes": 60,
      "upperBoundRunsThreeModels": 39,
      "upperBoundDeviceHoursThreeModels": 39
    },
    "seedList": [
      100,
      101,
      102,
      103,
      104
    ],
    "seedNote": "Example final-training seeds only. Publish separate split/search/initialisation seeds and all runs."
  },
  "compute": {
    "hardware": null,
    "softwareLock": null,
    "precision": "fp32",
    "trainingCapMinutes": 60,
    "latencyProtocol": {
      "warmup": 20,
      "measuredRepeats": 100,
      "report": [
        "median",
        "p95",
        "IQR"
      ],
      "scope": [
        "model-only",
        "full-scene end-to-end"
      ],
      "include": [
        "preprocessing",
        "patch extraction or tiling",
        "host-device transfers",
        "stitching"
      ],
      "acceleratorSynchronization": true,
      "coldCompileTimeSeparate": true
    },
    "genericAllocationIllustration": {
      "tokens": 49,
      "axis": "spectral",
      "bytesPerScalar": 4,
      "materializedScoresBytes": 38416,
      "oneTokenFeatureBytes": 12544,
      "oneStreamingStateBytes": 4096
    }
  },
  "outputsRequired": [
    "per-class precision/recall/F1",
    "macro-F1",
    "overall accuracy",
    "confusion matrix",
    "per-scene results",
    "paired seed results",
    "training/search time",
    "whole-scene latency",
    "peak allocated and reserved device memory",
    "host memory",
    "failed and OOM runs"
  ],
  "limitations": [
    "Matched input extent requires adapted models. Patch-restricting MambaHSI changes its whole-image method; this is not an exact paper reproduction.",
    "Scene-disjoint is a design intention here. Before running, verify scene IDs, provenance, duplicates and pretraining contamination."
  ],
  "measurements": {
    "accuracy": null,
    "latency": null,
    "peakMemory": null
  },
  "unfilledRequiredFields": [
    "datasetId",
    "sceneSplitManifestSha256",
    "labelBudgetPerClass",
    "trainValidationTestIds",
    "all model sourceRevision values",
    "hardware",
    "softwareLock"
  ]
}

Whole-image access is an information regime

For a strictly inductive, unseen-scene question, reserve evaluation scenes before fitting anything. Training may use complete training scenes; inference may use a complete held-out scene if that is how deployment works. Merely using a whole image at inference does not imply training leakage. What matters is when each scene becomes visible and which computations are allowed to depend on it.

For a transductive within-scene question, training can explicitly access unlabeled evaluation pixels. That can be a legitimate setting, but it is a different claim. Masking their labels in the loss does not remove their covariates from a whole-image forward pass, neighbourhood features, normalisation statistics or self-supervised objective. Report both label access and input access instead of using “no test labels” as a complete description.

Preprocessing needs the same audit. The inspected HybridSN notebook fits PCA to the reshaped cube; the inspected SpectralFormer demo obtains per-band minima and maxima from the image. Those are source-code facts, not accusations about every use of these methods. For a new inductive comparison, move learned preprocessing into the permitted training scope and document the deviation from the demonstration scripts. [4] [5]

Separate labelled centres and actual input footprints

Two disjoint lists of labelled pixel centres can still read overlapping source pixels. Liang and colleagues study this issue for spectral–spatial evaluation and explain why random same-image sampling can give misleading comparisons. For a claim about new regions or scenes, use an appropriate spatial or scene-level split and inspect the full processing footprint, including preprocessing filters. [7]

The lab’s toy geometry makes one part of that issue exact. Two 25 × 25 windows, centred eight columns apart at interior locations, share 17 × 25 = 425 of their 625 source pixels: 68%. Two 7 × 7 windows at those centres share none. Two full-image inputs share the entire 64 × 64 diagnostic grid. These are source-access counts, not model accuracy, effective receptive fields or proof that labels leaked.

Nonoverlapping footprints are helpful but do not establish statistical independence. Nearby fields, repeated specimens and overlapping acquisitions may remain related. A spatial buffer should account for every operation that mixes neighbours, and a scene holdout should match the intended geography, sensor, date or specimen transfer. Keep the richer split discussion separate from architecture claims, and publish the exact selected IDs.

Control spectral processing and pretraining separately

Changing the retained bands, band order, normalisation, PCA or wavelength resampling changes the learning problem. A PCA component index is not a wavelength. If a baseline receives projected components while another receives ordered bands, record that difference rather than describing both inputs only as “30 channels.” In a controlled track, begin with a shared, training-fitted representation that each adapted model can accept, then add preprocessing ablations.

Pretraining can add both information and computation. A from-scratch model and a pretrained model may be a reasonable deployment comparison, but their difference does not isolate architecture. Record corpus identity, scene overlap checks, available wavelengths, objective, checkpoint, frozen versus fine-tuned layers and adaptation budget. Shared corpus access still does not make objectives or realised training cost identical.

Keep a from-scratch or common-corpus track when the scientific claim needs it. A pretrained checkpoint whose data provenance cannot be audited should carry that uncertainty into the conclusion. The existing foundation-model guide covers broader transfer questions; here, the practical requirement is to account for these resources in the comparison, including any evaluation data used before supervised training.

Budget tuning, then repeat the selected experiment

Choose the selection rule before viewing test results. For example, select a configuration and checkpoint using validation macro-F1, with a fixed tie rule and training cap, then evaluate the locked choice on the test set. Do not select the best test epoch, best test seed or most favourable scene after the fact. Use the same class definitions and label budgets across candidates.

The builder proposes eight one-seed search trials and five final-training seeds per model. With a one-hour cap per run, three models have an upper budget of 39 single-device hours. These are editable planning choices, not a universal recommendation, an executed workload or a cloud price. Equal trial counts and equal time answer different resource questions; record both, along with unsuccessful and out-of-memory trials.

Split seeds and training seeds measure different variation. Reuse paired split manifests across methods, list all training runs, and report per-scene results plus paired differences. Five seeds is only an example; it cannot manufacture five independent scenes. PyTorch also warns that identical seeds do not guarantee identical results across releases, devices or platforms. Pin the software environment and record deterministic-kernel settings and any performance cost. [9]

Read complexity with the right sequence length

“Linear” and “quadratic” describe scaling under stated assumptions, not measured runtime for every input. Selective state-space sequence processing in Mamba motivates linear scaling with sequence length; it does not establish a universal HSI accuracy or latency advantage. Width, state size, number of passes, projections and implementation constants still matter. [8]

First identify the tokens. For the inspected SpectralFormer implementation, attention runs over band-derived tokens plus a class token. In MambaHSI’s spatial branch, sequence length follows the number of spatial pixels at that stage. Comparing one formula with L equal to bands against another with L equal to scene pixels is not a same-workload comparison. Likewise, spatial tiling, spectral grouping and patch extraction change the actual work. [3] [5] [6]

The resource panel intentionally computes only three named components: hL²s bytes for one materialized attention-score tensor, Lds bytes for one feature tensor, and dns bytes for one hypothetical streaming-state buffer. They are not interchangeable totals. FlashAttention uses an IO-aware exact-attention implementation that avoids the naive full-score storage pattern, so the first figure is not an unavoidable memory requirement of every Transformer. [10]

Measure the work the application actually needs

A useful latency report defines a workload: one isolated centre prediction, a batch of patches, or a complete output map for a specified scene. A whole-image call can produce thousands of predictions while a patch call produces a few. Comparing milliseconds per call without the output count is therefore misleading. Report both model-only inference and a full-scene end-to-end path.

For that end-to-end path, include preprocessing, patch extraction or tiling, data transfers and stitching. Disclose whether data are cached, whether compilation is already complete, and how boundaries are handled. Use the same device, precision policy, software stack where feasible, and stated CPU thread count. Synchronise asynchronous accelerator work before measuring completion, perform warmups, and report repeated timings with median and dispersion. PyTorch’s benchmark utilities explicitly handle these measurement issues. [11]

Peak device memory also needs a definition: allocated versus reserved memory, training versus inference, batch size, input shape and whether optimiser state already exists. Host memory and preprocessing buffers may be the actual bottleneck. MAC or FLOP counts require a convention and coverage of custom operators. Unsupported scan operators must not silently count as free. The lab has no timing engine and does not claim to estimate these quantities.

What the interactive lab really computes

With its defaults, a 25 × 25 × 48 FP32 patch contains 30,000 scalar values and needs 120,000 bytes. Sixteen patches need 1,920,000 bytes, about 1.83 MiB, for that raw input buffer alone. In native-inspired mode, a 128 × 128 × 48 FP32 full-scene input needs 3,145,728 bytes, exactly 3 MiB, and offers 16,384 prediction targets. Neither number includes model activations or weights.

For a naive dense map assembled from one same-size patch per output pixel, the logical input volume is H × W × p² × B values. The default is 491,520,000 scalar appearances. This counts repeated values and zero-padding positions. It is not a statement about actual disk reads, simultaneous RAM allocation or convolutional work; streaming and feature reuse can change those costs.

The generic spectral example uses L = 48 + 1 = 49, four heads and four bytes per scalar. Its single score tensor has 38,416 bytes. Doubling L would quadruple that component; doubling precision bytes doubles it. The downloaded Python script independently reproduces these counts and enumerates the overlap geometry on a CPU, without external dependencies, model downloads or GPU jobs. Its purpose is to make assumptions checkable.

Write a conclusion the evidence can support

A completed experiment should show per-class precision, recall and F1, macro-F1, overall accuracy, confusion matrices and scene-level behaviour. Where fine structures matter, add a defined boundary or small-object assessment. A high aggregate score can hide the classes or regions that motivated the application. Link each result to its configuration, split hash and checkpoint.

Report the chosen model’s useful trade-offs: accuracy under the stated information regime, training and search cost, memory, full-scene latency, deployment constraints and sensitivity to context. Include failed runs and adaptations. If results differ between matched and native tracks, discuss which information or implementation choices may explain the difference; a targeted ablation is more informative than a family-wide verdict.

This article stops before that empirical conclusion. No HybridSN, SpectralFormer or MambaHSI training was run for this deliverable, and no head-to-head accuracy, runtime or peak-memory measurements are presented. The exported protocol remains marked unexecuted, its measurement fields stay empty, and required dataset, revision, label-budget and hardware fields must be completed before it can describe a real study.

Source versions and references

  1. Roy, S. K. et al. HybridSN: Exploring 3D-2D CNN Feature Hierarchy for Hyperspectral Image Classification. arXiv:1902.06701v3 (3 July 2019); DOI 10.1109/LGRS.2019.2918719.
  2. Hong, D. et al. SpectralFormer: Rethinking Hyperspectral Image Classification with Transformers. arXiv:2107.02988v2 (20 November 2021); TGRS 60 (2022), DOI 10.1109/TGRS.2021.3130716.
  3. Li, Y. et al. MambaHSI: Spatial-Spectral Mamba for Hyperspectral Image Classification. TGRS 62 (2024); arXiv:2501.04944v1 (9 January 2025); DOI 10.1109/TGRS.2024.3430985.
  4. Paper-linked HybridSN repository, revision 8e9fd37 (23 December 2023). README and Hybrid-Spectral-Net.ipynb inspected on 2 October 2026.
  5. Official SpectralFormer repository, revision fd08eb5 (30 November 2024). README, demo.py and vit_pytorch.py inspected on 2 October 2026.
  6. Official MambaHSI repository, revision a705284 (9 April 2026). README, model/MambaHSI.py and train_MambaHSI.py inspected on 2 October 2026.
  7. Liang, J. et al. On the Sampling Strategy for Evaluation of Spectral-Spatial Methods in Hyperspectral Image Classification. TGRS 55(2), 862–880 (2017); author preprint arXiv:1605.05829.
  8. Gu, A. & Dao, T. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv:2312.00752v2 (2024).
  9. PyTorch 2.9 documentation: Reproducibility. Versioned documentation, accessed 2 October 2026.
  10. Dao, T. et al. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. arXiv:2205.14135v2 (2022).
  11. PyTorch 2.9 documentation: Benchmark Utils. Versioned documentation, accessed 2 October 2026.

Questions

Is Mamba always better than a CNN or Transformer for HSI?

No universal winner is established here. Results depend on the concrete implementation, data, information access, tuning and deployment workload; this article does not run a head-to-head benchmark.

Is whole-image inference automatically leakage?

No. A model trained on separate scenes can use the full held-out scene at inference. Evaluation covariates seen during training or preprocessing fitting must be separately disclosed.

Can the published accuracy tables be combined into a leaderboard?

Only after establishing compatible datasets, labels, splits, preprocessing, context and evaluation procedures. The article avoids cross-paper numerical ranking.

Does a matched patch experiment reproduce MambaHSI?

Restricting the whole-image design to isolated patches changes it. Document the adaptation and report it separately from a native-pipeline reproduction.

Does the resource calculator predict GPU memory or latency?

No. It counts raw inputs and explicitly named tensor components, omitting actual model state, activations, kernel workspace and runtime. Its streaming-state illustration is not a Mamba training-memory estimate.

Does exporting the protocol start an experiment?

No. It downloads an unexecuted plan with null measurement and provenance fields. The CPU script performs only arithmetic and toy geometry; it does not train or download models.