A GPU board, memory modules, solid-state storage and a compact spectrometer rest on a modern lab workstation.
Conceptual artwork. Diagrams and examples below explain the technical details.

Start with the workload and the data representation

Hyperspectral analysis can mean reading a few spectra, computing a scene-wide transform or training a neural network on patches. These tasks put different pressures on a computer. Begin by describing the operation, the cube dimensions, the stored data type and the working data type. A file’s size on disk is useful context, but it does not tell you the largest amount of memory an algorithm will need.

The discussion here is a workload guide using public software documentation and synthetic calculations. It does not rank products or specify a shopping list. The aim is to make a small experiment measurable, then use those measurements to decide how to process a larger cube.

Track representation and memory at each boundary in an HSI pipeline.Cube and metadata to GDAL or HSI reader: blocks or windows; GDAL or HSI reader to Array in system RAM: shape and data type; Array in system RAM to CPU analysis: numerical operations; Array in system RAM to GPU tensor workload: device transfer; CPU analysis to Results and provenance: write; GPU tensor workload to Results and provenance: copy and writeConceptual relationshipsCube and metadataGDAL or HSI readerArray in systemRAMCPU analysisGPU tensorworkloadResults andprovenanceblocks or windowsshape and data typenumericaloperationsdevice transferwritecopy and write
Conceptual illustration. Track representation and memory at each boundary in an HSI pipeline.

Calculate the cube before loading it

For a dense cube, element storage is rows multiplied by columns, bands and bytes per value. A 1,024 × 1,024 × 224 cube represented as float32 has 939,524,096 bytes of elements, or 0.875 GiB. A 4,096 × 4,096 × 224 cube has 14 GiB of float32 elements. The same dimensions require 7 GiB as uint16 or 28 GiB as float64. GiB uses powers of 1,024; decimal GB uses powers of 1,000.

These are calculations of uncompressed element storage. They exclude array-object overhead, masks, caches, model state and temporary arrays. NumPy’s nbytes reports bytes occupied by an array’s elements and excludes its non-element attributes. Treat it as one measurement in a memory budget, rather than the memory use of the entire Python process.

If an operation keeps an input cube, an output cube and one equally sized temporary array, the illustrative float32 budget becomes 42 GiB for the larger cube. That is an explicitly assumed three-buffer example. Actual peak memory depends on the algorithm, implementation, conversions and when buffers are released. Check the peak on a representative run before scaling.

Sources: [2]

Calculated element storage for dense synthetic arrays; excludes overhead and algorithm state.
Shape: rows × columns × bandsRepresentationBytesBinary capacity
1,024 × 1,024 × 224float32939,524,0960.875 GiB
4,096 × 4,096 × 224uint167,516,192,7687 GiB
4,096 × 4,096 × 224float3215,032,385,53614 GiB
4,096 × 4,096 × 224float6430,064,771,07228 GiB
256 × 256 × 224float3258,720,25656 MiB

CPU work covers more than model training

The CPU usually participates in file reading, decompression, metadata checks and preparation of arrays. Many spectral operations also run there, including reductions, matrix calculations and classical analysis. The important distinction is whether an operation spends time waiting for data, moving data through memory or performing arithmetic. Increasing parallel workers can help some workloads and duplicate memory or compete for storage in others.

Measure stages separately on a representative subset: opening the file, reading values, converting the data type and running the calculation. Keep the same input and parameters when comparing runs. Record the number of workers and numerical-library threads, because their interaction can affect timings. This measurement plan is more informative than assuming that every HSI pipeline benefits from the same CPU configuration.

Estimate raw cube memory

Enable JavaScript to explore the calculated chart.

Raw uncompressed float32 cube: 1024 × 1024 pixels. Model activations, optimiser state and I/O buffers are additional.

RAM holds working state; storage supplies it

System RAM must accommodate the arrays that are live at the same time, together with the application and library overhead. Storage capacity must accommodate source files, intermediate products, checkpoints and final outputs. Storage access speed matters when a pipeline repeatedly reads patches or writes large results. An SSD does not remove the need to understand the access pattern, and unused RAM does not make an inefficient read pattern disappear.

Inspect how often each stage reads the same values. A small preview may require three bands, while a spectral transform may require every band for every selected pixel. Arrange experiments so the input read, computation and output write can be timed independently. Keep a record of whether the data was already in operating-system caches; otherwise, apparently faster repeats may be measuring a different situation.

GPU memory is a separate budget

A GPU can be useful for tensor operations and neural-network workloads that are implemented for it. Training needs room for more than input patches: model parameters, activations, gradients and optimiser state also consume device memory. Inference has a different budget. Copying tensors between system memory and the device has a cost, so a fast individual kernel does not guarantee a fast complete pipeline.

PyTorch’s CUDA documentation distinguishes memory occupied by tensors from memory reserved by its caching allocator. Its allocated and reserved memory measurements therefore answer different questions. The documentation also explains that empty_cache releases unused cached memory; it does not free storage occupied by live tensors. Record relevant peaks and the batch, patch and data type used.

A useful experiment holds the model and patch dimensions fixed while changing batch size, then measures throughput and peak device memory. Verify numerical behaviour when changing precision. The experiment should include data loading and transfer as well as model execution if the intended claim concerns total processing time.

Sources: [5]

Give each software layer a clear responsibility

Python provides the surrounding scripts, configuration and run records. Its venv module creates environments with their own installed packages, isolated by default from the base environment’s packages. Recreating an environment still requires the interpreter and dependency information. Keep that information with the project; an environment directory alone is not a portable research record.

NumPy provides multidimensional arrays and numerical operations. GDAL’s raster data model represents datasets, raster bands, georeferencing and metadata. Rasterio exposes raster operations in Python and documents windowed reading and writing. Together these layers let a pipeline preserve geospatial context while working with array values. Check axis order and units at the boundary between readers and analysis code.

Spectral Python provides interfaces intended for hyperspectral data, including ENVI-backed images. Its documented SpyFile interface initially reads metadata and obtains image values on demand. Its loaded ImageArray uses rows, columns and bands with 32-bit floating-point elements. That representation can differ from the source file’s layout and data type, so loading the image may change the memory requirement.

Sources: [1], [2], [4], [3], [6]

Windowed reads help only when the operation permits them

A 256 × 256 × 224 float32 tile contains 56 MiB of elements. Non-overlapping tiles of that size cover the illustrative 4,096 × 4,096 scene in 256 windows. This makes local operations possible without holding the full float32 cube. If a spatial filter needs neighbouring pixels, the read window must include a suitable border, and the output must handle those overlaps deliberately.

Rasterio notes that a windowed request reads the underlying storage blocks needed to fill the window. A one-pixel request can therefore cause a much larger read. Aligning windows with the dataset’s block structure can improve efficiency. Windowing also does not automatically turn a global algorithm into a local one: a scene-wide statistic or transform may need a separate pass, accumulated state or a method designed for streaming.

Sources: [3]

Memory mapping and loading solve different access problems

Spectral Python documents memory mapping as an alternative to loading the entire image. It also notes that its ordinary on-demand image reads are not cached by SpyFile, so iterative access can carry a penalty. The right choice depends on which values are revisited and how often. Inspecting a few spectra and repeatedly fitting an algorithm over the whole cube are different cases.

File-backed access does not guarantee that later calculations remain small. Conversions and expressions may create arrays. Check intermediate shapes and data types, then measure the process peak. Compare access methods using the same numerical task and output checks.

Sources: [6]

Use a measured scaling plan

Begin with a small public or synthetic case and verify the output. Then increase spatial size, band count or batch size one variable at a time. Record wall time, peak system memory, peak device memory where relevant and the amount of data read and written. Keep the data type explicit. A calculation that fits in memory is a capacity result; a calculation that finishes quickly is a separate performance result.

The standard-library example below calculates capacities without allocating an image. It makes the assumptions visible and produces the exact figures used in this article. It cannot predict a particular algorithm’s runtime or peak memory. Use it to reject an impossible starting plan, then profile the real operation with its actual readers, transforms and outputs.

Run the example

Prerequisite: Python 3. Examples use synthetic inputs to explain the calculation. Save the snippet as example.py and run python3 example.py.

from math import ceil

rows, columns, bands = 4096, 4096, 224
bytes_per_value = {"uint16": 2, "float32": 4, "float64": 8}
for name, itemsize in bytes_per_value.items():
    size = rows * columns * bands * itemsize
    print(f"{name}: {size / 1024**3:.1f} GiB")
float32_size = rows * columns * bands * bytes_per_value["float32"]
print(f"three float32 buffers: {3 * float32_size / 1024**3:.1f} GiB")
tile_rows, tile_columns = 256, 256
tile_size = tile_rows * tile_columns * bands * bytes_per_value["float32"]
tile_count = ceil(rows / tile_rows) * ceil(columns / tile_columns)
print(f"one float32 tile: {tile_size / 1024**2:.1f} MiB")
print(f"non-overlapping tiles: {tile_count}")

Verified output

uint16: 7.0 GiB
float32: 14.0 GiB
float64: 28.0 GiB
three float32 buffers: 42.0 GiB
one float32 tile: 56.0 MiB
non-overlapping tiles: 256

Frequently asked questions

Does a 14 GiB cube require only 14 GiB of RAM?

That figure covers float32 elements only. Temporary arrays, masks, caches and other live state can increase process memory substantially.

Are GB and GiB interchangeable?

No. GB is 1,000,000,000 bytes and GiB is 1,073,741,824 bytes. State which unit you use in capacity calculations.

Do I need a GPU to analyse HSI?

Many inspection and classical numerical tasks run on the CPU. GPU usefulness depends on the implemented workload, transfers and available device memory.

Can I process every algorithm tile by tile?

Local operations often allow windowing. Global statistics and transforms need suitable accumulated state, multiple passes or an algorithm designed for streaming.

Does memory mapping eliminate array copies?

No. Access can be file-backed while later conversions and computations create new arrays. Check intermediates and measure the process peak.

What should a timing report include?

Input shape and data type, operation, versions, workers, batch or window size, cache conditions and whether reading, transfer and writing are included.

References and further reading

  1. Python: venv
  2. NumPy: ndarray.nbytes
  3. Rasterio: Windowed reading and writing
  4. GDAL: Raster Data Model
  5. PyTorch: CUDA semantics
  6. Spectral Python: Reading HSI Data Files