HSI preprocessing playground: complete synthetic CPU example (v2)

REQUIREMENTS
Python 3.12 recommended. Tested on Python 3.12.14 with pinned requirements.
NumPy, SciPy and scikit-learn are the only numerical dependencies. No GPU,
credentials, private data or external dataset is needed. The notebook contains
all functions and does not require the companion script to execute.

QUICK START (after extracting the ZIP into a folder)
  python -m venv .venv
Activate the environment using your platform's normal command:
  macOS/Linux: source .venv/bin/activate
  Windows PowerShell: .venv\Scripts\Activate.ps1
Then:
  python -m pip install -r requirements.txt
  python hsi_preprocessing.py --self-test
  python hsi_preprocessing.py

REPRODUCE A RECIPE AND SAVE THE SPECTRUM
  python hsi_preprocessing.py --preset scatter --normalisation msc --operation derivative2 --window 11 --output example.csv
  python hsi_preprocessing.py --preset baseline --detrend 2 --normalisation snv --operation moving_average --window 5 --output baseline.csv
  python hsi_preprocessing.py --gain 1.25 --offset 0.08 --noise 0.02 --curve 0.05 --normalisation l2 --operation smooth --output custom.csv
Each requested CSV has a companion .metadata.json recording units and settings.
The playground also displays an exact CLI command for its current settings.
  python hsi_preprocessing.py --help

METHODS
--normalisation: none, minmax, zscore, snv, l2, msc, feature_zscore
--detrend: 0 (off), 1 (linear), 2 (quadratic)
--operation: none, smooth, derivative, derivative2, moving_average
--window: 5, 11, 21 bands (20, 50, 100 nm first-to-last spans)
--preset: clean, noise, scatter, baseline, mixed
--sample: 1, 2 or 3
Optional custom controls: --gain, --offset, --noise, --curve
--include-flagged disables the known synthetic bad-band mask.

Semantics: mask -> detrend -> normalise/scatter-correct -> filter.
minmax/zscore/snv/l2 act within each spectrum. SNV uses ddof=1; row zscore uses
ddof=0. MSC uses a fixed training-mean reference. feature_zscore uses means and
population scales per band learned across training observations. The fit method
learns those statistics inside each training fold. Never fit them on test rows.
Detrending and filters act separately in each contiguous valid run. Moving-average
windows shrink at segment edges; SG uses degree 2 and polynomial edge fits.
First and second derivatives are per nm and per nm squared, respectively.

VALIDATION
  python -m unittest -v test_preprocessing.py
25 unittest checks; additionally the browser is compared with 630 Python recipe
fixtures. To regenerate those fixtures into a chosen directory:
  python hsi_preprocessing.py --export ./fixtures
The notebook has been executed cell by cell in order; it does not install packages.

LIMITS
All data are synthetic. Every preset retains an injected corrupt 1340-1390 nm
interval; it is not a universal bad-band recommendation. Quiet/clean means no
additional noise/gain/offset/curvature beyond the fixed underlying synthetic
sample and that flagged interval. This is a practical subset, not all methods.
A smoother plot or normalised range is not a prediction-accuracy improvement.

The combined functions deliberately require finite input data on a complete,
strictly increasing uniform wavelength grid before masking. They reject
compressed gapped/irregular coordinates instead of guessing a delta. Row
normalisation itself does not mathematically require a uniform grid. Establish
an explicit missing-data/resampling policy before adapting to a real product.
Zero/near-zero normalisation denominators, flat MSC references and non-positive
MSC gains raise errors rather than being silently repaired. These numerical
guards are not instrument-specific quality thresholds. Detrending and scaling
can remove useful information. Validate the whole recipe on independent units.

PRIMARY SOURCES
https://docs.scipy.org/doc/scipy/reference/generated/scipy.signal.savgol_filter.html
https://doi.org/10.1366/0003702894202201
https://doi.org/10.1366/0003702854248656
https://scikit-learn.org/stable/common_pitfalls.html
https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.Normalizer.html
https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.MinMaxScaler.html
