Spectral methods lab · original synthetic data

48 bands. Different questions.

Fit a projection, inspect a spectrum, and see what each objective keeps. All 120 samples are synthetic. No GPU, API key, or upload is needed.

PCA, LDA and the autoencoder fit only the 80 training spectra. t-SNE and UMAP presets fit all 120 spectra together and are exploratory.

Enable JavaScript to compute the projection. The downloadable Python code can reproduce every method.

Each mark is one synthetic spectrum. Coordinates are method-specific; separate plots do not share a distance scale.
● A · 550 nm dip■ B · 670 nm dip▲ C · 810 nm dip◆ D · 930 nm dip

Filled = original training split (80). Hollow = original held-out split (40). For t-SNE/UMAP, both groups participate in fitting.

Displayed embedding
Awaiting fit

Loading the small synthetic dataset. Controls stay disabled until it is ready.

Inspect the original spectrum

Wavelengths are a synthetic 450–1000 nm grid. Values are illustrative unitless reflectance-like numbers, not calibrated measurements.

What did the method keep?

Explained variance measures training variance, not classification performance.

Read the selected spectrum as numbers
Original and reconstructed values for the currently selected sample; all 48 synthetic bands
Wavelength (nm)OriginalReconstructed
Read the projection as numbers
Coordinates shown in the current two-dimensional display
SampleClassOriginal splitAxis 1Axis 2
Data, fitting scope and honest comparison

This fixed generator creates 30 independent spectra in each of four toy classes. A broad brightness nuisance, smooth latent variation and small noise accompany a class-specific absorption-like dip. It is not a sensor simulator or a field HSI benchmark. Samples 1–20 within each class train the learned methods; the last 10 are held out. A real HSI study needs splits by appropriate specimen, field, scene or acquisition groups.

The mean and optional per-band standard deviation always use only the 80 training spectra. PCA, Fisher LDA and the autoencoder fit those 80 rows. LDA receives only training labels. All 120 points participate in each stored t-SNE/UMAP fit; their hollow marks are original split identifiers, not out-of-sample projections. Class colours are annotations for every unsupervised method.

Neighbour overlap is the average fraction of five nearest neighbours in preprocessed 48-band space also found in the displayed 2D coordinates, over all 120 points. It is a descriptive geometry diagnostic, not accuracy, a test score or evidence of physical classes. For one retained component, axis 2 is zero. With more than two retained dimensions, the plot and overlap still use only the first two. Raw-space reconstruction MSE uses all 48 bands and all retained dimensions.

Changing settings after looking at held-out errors turns this toy holdout into a development set. Reserve a genuinely untouched test set for real conclusions. Rotation, reflection, visual cluster area and distances between t-SNE/UMAP islands are not reliable physical explanations.

t-SNE presets: scikit-learn exact solver, random initialisation, 750 iterations, automatic learning rate. UMAP presets: umap-learn, Euclidean distance, spectral initialisation, 300 epochs, one CPU thread. Seeds 7 and 29; all 36 actual outputs and version metadata ship with the source. The browser never pretends to fit these stored runs.

Runnable source is included in this folder.Python notebook