{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# Build a reproducible HSI pipeline in Python\n",
        "\n",
        "A CPU-only synthetic teaching example. Extract the full companion ZIP and open\n",
        "this notebook with that folder as the working directory: it imports the adjacent\n",
        "`hsi_pipeline.py`, so the script and notebook share one implementation.\n",
        "\n",
        "Tested: CPython 3.12.14 and NumPy 2.3.5. The notebook installs nothing, downloads\n",
        "nothing and uses no GPU. All spectra, labels and reported scores are generated;\n",
        "none demonstrate performance on measured HSI data.\n",
        "\n",
        "## 1. Generate and validate the inputs"
      ],
      "id": "hsi-pipeline-00"
    },
    {
      "cell_type": "code",
      "metadata": {},
      "source": [
        "from pathlib import Path\n",
        "import json\n",
        "import tempfile\n",
        "import numpy as np\n",
        "from hsi_pipeline import (Config, synthetic_cube, fit_model, audit_train_only,\n",
        "                          validate_inputs, validate_masks, run, verify_run,\n",
        "                          load_model, predict)\n",
        "cfg = Config()\n",
        "cube, labels, wavelengths, keep, masks = synthetic_cube(cfg)\n",
        "print(\"Cube (row, column, band):\", cube.shape)\n",
        "print(\"Wavelengths (nm):\", wavelengths[0], \"to\", wavelengths[-1])\n",
        "print(\"Retained bands:\", int(keep.sum()), \"of\", len(keep))\n",
        "print(\"Split counts:\", {k: int(v.sum()) for k, v in masks.items()})\n",
        "print(\"Unused:\", int((~np.logical_or.reduce(list(masks.values()))).sum()))"
      ],
      "execution_count": 1,
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Cube (row, column, band): (60, 90, 41)\n",
            "Wavelengths (nm): 900.0 to 1700.0\n",
            "Retained bands: 38 of 41\n",
            "Split counts: {'train': 1800, 'validation': 1320, 'test': 1800}\n",
            "Unused: 480\n"
          ]
        }
      ],
      "id": "hsi-pipeline-01"
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "## 2. Fit only training pixels\n",
        "\n",
        "The baseline uses per-band training means and population standard deviations\n",
        "(ddof=0), followed by class centroids. A saved training fit-ID vector makes the\n",
        "membership explicit. Four unused columns separate the stripe masks; this is not\n",
        "a universal independence or receptive-field guarantee."
      ],
      "id": "hsi-pipeline-02"
    },
    {
      "cell_type": "code",
      "metadata": {},
      "source": [
        "model = fit_model(cube, labels, wavelengths, keep, masks)\n",
        "audit_train_only(model, cube, masks[\"train\"])\n",
        "assert len(model[\"fit_pixel_ids\"]) == int(masks[\"train\"].sum())\n",
        "print(\"Fitted pixels:\", len(model[\"fit_pixel_ids\"]))\n",
        "print(\"Centroid matrix:\", model[\"centroids\"].shape)\n",
        "print(\"Train-only fit audit passed\")"
      ],
      "execution_count": 2,
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Fitted pixels: 1800\n",
            "Centroid matrix: (3, 38)\n",
            "Train-only fit audit passed\n"
          ]
        }
      ],
      "id": "hsi-pipeline-03"
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "## 3. Watch deliberate mistakes fail\n",
        "\n",
        "These are expected validation errors. A successful test is one that detects the\n",
        "mistake. Do not remove an assertion to make corrupted data pass."
      ],
      "id": "hsi-pipeline-04"
    },
    {
      "cell_type": "code",
      "metadata": {},
      "source": [
        "def expect_failure(label, operation):\n",
        "    try:\n",
        "        operation()\n",
        "    except ValueError as error:\n",
        "        print(label + \":\", error)\n",
        "    else:\n",
        "        raise AssertionError(label + \" was not detected\")\n",
        "\n",
        "expect_failure(\"Reversed wavelengths\", lambda: validate_inputs(\n",
        "    cube, labels, wavelengths[::-1], keep, masks))\n",
        "wrong_masks = {k: v.copy() for k, v in masks.items()}\n",
        "wrong_masks[\"test\"][0, 0] = True\n",
        "expect_failure(\"Split overlap\", lambda: validate_masks(wrong_masks, labels.shape))\n",
        "leaky = {k: v.copy() for k, v in model.items()}\n",
        "leaky[\"mean\"] = cube.reshape(-1, cube.shape[-1])[:, keep].mean(axis=0)\n",
        "expect_failure(\"Global-fit scaler\", lambda: audit_train_only(leaky, cube, masks[\"train\"]))"
      ],
      "execution_count": 3,
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Reversed wavelengths: Wavelengths must be positive and strictly increasing\n",
            "Split overlap: Split overlap detected\n",
            "Global-fit scaler: Preprocessing leakage or mismatched training statistics\n"
          ]
        }
      ],
      "id": "hsi-pipeline-05"
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "## 4. Perturb the holdout without changing the fitted state\n",
        "\n",
        "Changing held-out spectra and labels must not change the trained state. This\n",
        "checks the implementation boundary, not whether a researcher inspected test\n",
        "results while choosing the overall recipe."
      ],
      "id": "hsi-pipeline-06"
    },
    {
      "cell_type": "code",
      "metadata": {},
      "source": [
        "altered_cube, altered_labels = cube.copy(), labels.copy()\n",
        "heldout = masks[\"validation\"] | masks[\"test\"]\n",
        "altered_cube[heldout] += 500\n",
        "altered_labels[heldout] = (altered_labels[heldout] + 1) % 3\n",
        "altered_model = fit_model(altered_cube, altered_labels, wavelengths, keep, masks)\n",
        "for name in model:\n",
        "    np.testing.assert_array_equal(model[name], altered_model[name], err_msg=name)\n",
        "print(\"Every fitted array is unchanged by holdout perturbation\")"
      ],
      "execution_count": 4,
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Every fitted array is unchanged by holdout perturbation\n"
          ]
        }
      ],
      "id": "hsi-pipeline-07"
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "## 5. Save a complete run, then verify it\n",
        "\n",
        "Use a fresh directory for every run. The manifest is written last and lists data,\n",
        "masks, config, environment, model state, predictions, metrics, figures and exact\n",
        "executed source. It checks integrity relative to itself; it is not an author\n",
        "signature or scientific-validity certificate."
      ],
      "id": "hsi-pipeline-08"
    },
    {
      "cell_type": "code",
      "metadata": {},
      "source": [
        "temporary = tempfile.TemporaryDirectory(prefix=\"hsi-notebook-\")\n",
        "run_dir = Path(temporary.name) / \"run\"\n",
        "manifest = run(run_dir, cfg)\n",
        "print(verify_run(run_dir))\n",
        "result = json.loads((run_dir / \"evaluation/test_metrics.json\").read_text())\n",
        "print(\"Test confusion matrix (reference rows, prediction columns):\")\n",
        "print(np.array(result[\"confusion_matrix\"]))\n",
        "print(\"OA:\", result[\"overall_accuracy\"])\n",
        "print(\"Macro-F1:\", result[\"macro_f1\"])\n",
        "assert sum(map(sum, result[\"confusion_matrix\"])) == int(masks[\"test\"].sum())\n",
        "print(\"Counts agree; these are synthetic regression fixtures\")"
      ],
      "execution_count": 5,
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "{'verified_artifacts': 32, 'status': 'checksums match', 'caution': 'Integrity relative to this manifest, not authenticity or scientific validity'}\n",
            "Test confusion matrix (reference rows, prediction columns):\n",
            "[[591  45   0]\n",
            " [  0 582   0]\n",
            " [  0   0 582]]\n",
            "OA: 0.975\n",
            "Macro-F1: 0.9753681132338755\n",
            "Counts agree; these are synthetic regression fixtures\n"
          ]
        }
      ],
      "id": "hsi-pipeline-09"
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "## 6. Reload the model and compare two complete reruns\n",
        "\n",
        "This is a same-environment check. A seed does not guarantee identical numerical\n",
        "results across different builds, versions, operating systems or devices."
      ],
      "id": "hsi-pipeline-10"
    },
    {
      "cell_type": "code",
      "metadata": {},
      "source": [
        "restored = load_model(run_dir / \"model\")\n",
        "prediction = predict(restored, cube[masks[\"test\"]], wavelengths)\n",
        "stored = np.load(run_dir / \"evaluation/test_predictions.npy\", allow_pickle=False)\n",
        "np.testing.assert_array_equal(prediction, stored)\n",
        "second_dir = Path(temporary.name) / \"rerun\"\n",
        "second_manifest = run(second_dir, cfg)\n",
        "assert manifest == second_manifest\n",
        "print(\"Reloaded predictions agree\")\n",
        "print(\"Two runs have identical SHA-256 artifact records in this environment\")"
      ],
      "execution_count": 6,
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Reloaded predictions agree\n",
            "Two runs have identical SHA-256 artifact records in this environment\n"
          ]
        }
      ],
      "id": "hsi-pipeline-11"
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "## 7. Inspect the saved figures\n",
        "\n",
        "The bundle's `example-run/figures/spatial-receipt.svg` shows the reference map,\n",
        "fixed split masks and test-only predictions. Pixels outside the test stripe are\n",
        "not evaluated. The confusion figure includes every numeric count.\n",
        "\n",
        "![Reference, fixed masks and held-out map. All data are synthetic.](example-run/figures/spatial-receipt.svg)\n",
        "\n",
        "![Synthetic confusion matrix with rows 591/45/0, 0/582/0 and 0/0/582.](example-run/figures/confusion-matrix.svg)\n",
        "\n",
        "## 8. Clean up this notebook's temporary runs\n",
        "\n",
        "The supplied `example-run/` stays untouched. To retain a new run outside the\n",
        "notebook, use the command-line interface with a fresh output directory."
      ],
      "id": "hsi-pipeline-12"
    },
    {
      "cell_type": "code",
      "metadata": {},
      "source": [
        "temporary.cleanup()\n",
        "print(\"Notebook temporary runs removed; supplied example-run is unchanged\")"
      ],
      "execution_count": 7,
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Notebook temporary runs removed; supplied example-run is unchanged\n"
          ]
        }
      ],
      "id": "hsi-pipeline-13"
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "## What to change before real research\n",
        "\n",
        "Replace the synthetic reader with a validated acquisition reader. Preserve units,\n",
        "no-data rules, calibration provenance and annotation meaning. Choose independent\n",
        "splits for the claim you need; refit learned transformations inside each fold.\n",
        "Keep final-test feedback out of model development. Record richer hardware and\n",
        "backend details when adding numerical libraries or accelerators.\n",
        "\n",
        "The package contains 18 unittest cases: run `python -m unittest -v test_pipeline.py`.\n",
        "\n",
        "Primary sources:\n",
        "- [NumPy random compatibility](https://numpy.org/doc/stable/reference/random/compatibility.html)\n",
        "- [scikit-learn leakage guidance](https://scikit-learn.org/stable/common_pitfalls.html)\n",
        "- [PyTorch reproducibility](https://docs.pytorch.org/docs/stable/notes/randomness.html)\n",
        "- [Barker and colleagues, FAIR4RS (2022)](https://doi.org/10.1038/s41597-022-01710-x)\n",
        "\n",
        "The saved outputs below were verified by sequential Python cell execution, not a\n",
        "Jupyter-kernel run. Notebook presentation depends on the reader's notebook app."
      ],
      "id": "hsi-pipeline-14"
    }
  ],
  "nbformat": 4,
  "nbformat_minor": 5,
  "metadata": {
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "name": "python",
      "version": "3.12.14"
    }
  }
}
