Conceptual NIR spectrometer with a prepared chickpea flour sample and whole chickpeas
AI-generated editorial illustration of prepared flour and a generic NIR instrument. It is not a photograph of the study setup or an exact instrument reconstruction.

The measurement is a spectrum, and the target comes from a laboratory

What can one gram of flour tell us about a chickpea variety? In our published study, the answer came from pairing a visible–near-infrared reflectance spectrum with laboratory measurements of composition. The instrument records how a sample interacts with light; a calibrated regression model connects that signal to an independently measured property. It does not count protein molecules directly.

The study, Unlocking chickpea flour potential: AI-powered prediction for quality assessment and compositional characterisation, used flour from 136 chickpea varieties. Each observation contained 1,921 spectral features and six targets: protein, starch, soluble sugars, insoluble fibres, total lipids and moisture by mass. This is a spectral-only regression problem. There is no spatial image or per-pixel composition ground truth in this experiment.

This article follows that measurement chain and explains the Deep Learning Grid Explorer Framework, or DLGX, described in my confirmation presentation. The framework organised exploratory model experiments; the peer-reviewed chickpea paper provides the published evidence for the application. The demonstration below uses synthetic spectra rather than the study data.

Sources: [1]

Prepared flour provides both the spectrum and laboratory reference targets used to develop a calibration.Prepared flour sample to NIR reflectance spectrum; Prepared flour sample to Laboratory reference assays; NIR reflectance spectrum to Paired data and calibration; Laboratory reference assays to Paired data and calibration; Paired data and calibration to Independent prediction checksConceptual relationshipsPrepared floursampleNIR reflectancespectrumLaboratoryreference assaysPaired data andcalibrationIndependentprediction checks
Conceptual illustration. Prepared flour provides both the spectrum and laboratory reference targets used to develop a calibration.

Prepare the flour before interpreting the light

The varieties were grown under common glasshouse conditions in Perth, Western Australia. The grain was soaked at 4 °C to soften the seed coat, decorticated by hand, dried at 40 °C and milled using a TissueLyzer II with a 2 cm ball bearing at 25 Hz for two minutes. Flour was stored in airtight containers at room temperature. These reported preparation steps define the material on which the models were evaluated.

For spectroscopy, the paper reports 1 g of milled flour per sample. Particle size, packing, surface presentation and temperature can change the optical response even when the chemical composition is similar. Consistent sample preparation reduces avoidable measurement variation. Repacking and repeated measurement would be useful checks of presentation sensitivity, but the paper does not specify a repeat-scan or repacking protocol that can be reconstructed here.

“Non-destructive” describes the spectroscopic reading of the prepared sample. Milling, decortication and the chemical reference assays are separate preparation or analysis operations. A model developed on this prepared flour should not automatically be applied to intact seeds, different milling conditions or a production line.

Sources: [1], [2]

What the measurement chain provides, and what still requires validation
StageOutputInterpretation limit
Prepared flour1 g sample used in the studyPreparation and packing affect what light samples
Spectrometer1,921 reflectance features; 680–2,600 nm1 nm sampling is not proof of 1 nm optical resolution
Wet laboratorySix sample-level composition targetsAssay definitions and uncertainty remain relevant
DLGX explorationModel, preprocessing and prediction comparisonsFramework organisation is not independent validation
Calibrated regressionPredicted composition in target unitsValid only within a demonstrated calibration domain

Capture diffuse reflectance, not a hyperspectral image

The reported instrument was the Unity Scientific SpectraStar 2600XT-R. It provided reflectance from 680 to 2,600 nm at 1 nm sampling. Including both endpoints gives (2600 − 680)/1 + 1 = 1,921 values. Sampling interval describes the wavelength grid; it does not establish that the optical spectral resolution is 1 nm.

In a diffuse-reflectance measurement, light entering a powder can undergo multiple scattering events and absorption before some of it returns to the collection optics. The detector therefore responds to both chemical absorption and physical presentation. The 3D lab separates illumination, powder and collected return light to make those relationships visible. It is an explanatory geometry, not a reconstruction of the manufacturer’s internal optics or the exact cup used in the study.

A bulk spectrum averages over an optical sampling region. The schematic beam footprint is not a pixel map, and the animated wavelength sweep is a deliberately slowed reading sequence. It does not reproduce the instrument’s timing, mechanical motion, detector design or reference routine.

Sources: [1], [2]

INTERACTIVE MEASUREMENT LAB · SYNTHETIC DATA

Follow light through a flour measurement

A rotatable schematic explains diffuse reflectance. Change the synthetic spectrum, then reveal its wavelength samples. The 3D geometry and playback timing do not reproduce the SpectraStar’s internal design.

Prepare a consistent flour sample. The paper reports 1 g per measurement.

Schematic labels: 1 illumination · 2 prepared flour · 3 returned-light collection. Gold traces show illumination; teal traces show selected scattered return paths. Colours indicate paths, not the visible colour of NIR radiation.

Enable JavaScript to explore the spectrum.

1,921 wavelength samples, one sample-level spectrum. Gaussian teaching bands and an additive offset generate the curve; no chemical concentrations or trained-model predictions are returned. SNV here is per-spectrum centring and division by sample standard deviation. It is not the paper’s selected PP-04 pipeline.

Reference the instrument and distinguish reflectance from absorbance

A general linear-response teaching model is R(λ) = Rref(λ) × [S(λ) − D(λ)]/[W(λ) − D(λ)]. S is the sample signal, W an illuminated reference, D the blocked-light detector signal, and Rref the known reference reflectance. Comparable exposure, illumination and geometry matter. Saturated signals and reference bands close to the dark level make that calculation unreliable.

This equation explains why lamp output and detector response cannot simply be interpreted as composition. The paper reports reflectance measurements but does not give enough detail to reproduce its dark/reference acquisition, scan averaging or cup configuration. Those operations in the lab are general teaching steps, not undocumented claims about the experiment.

For positive fractional reflectance, pseudo-absorbance is A(λ) = −log10 R(λ). A decrease in reflectance becomes an increase in pseudo-absorbance. In scattering flour this is a useful representation rather than a direct Beer–Lambert concentration measurement with one known path length. The lab lets you change a synthetic absorption strength and a scattering offset independently, then compare reflectance, pseudo-absorbance and per-spectrum standard normal variate (SNV).

Sources: [1], [2]

Read broad, overlapping molecular information

NIR bands commonly arise from overtones and combinations of molecular vibrations. O–H, N–H and C–H containing groups contribute to the signals of water, protein, carbohydrates and lipids. These features overlap; predicting composition usually requires information across several wavelengths and reference-labelled calibration samples.

Moisture-related O–H features commonly occur around 1,400–1,500 nm and 1,900–2,000 nm. The demonstration includes broad synthetic depressions around 1,450 and 1,940 nm to explain this behaviour. Their depths and shapes are invented teaching values. Other drawn bands are generic overlapping features; they are not unique chemical identifiers.

The absorption slider has arbitrary units. It cannot estimate a flour’s moisture percentage, protein content or the study’s regression outputs. Its purpose is to expose an identifiability problem: a chemical change and a physical scattering change can both alter a spectrum. Preprocessing and representative calibration data help a model distinguish useful variation from nuisance variation.

Sources: [2]

Pair spectra with six independently measured targets

The chemical reference analyses are what make supervised composition prediction possible. The paper describes ethanol extraction and an anthrone assay for soluble sugars, starch digestion with a Megazyme total starch assay, proteinase K treatment and dried residual mass for insoluble fibre, a Bradford assay for protein, and chloroform–methanol extraction followed by gravimetry for lipids. Moisture by mass is the sixth reported target; its full reference protocol should be taken from the associated methods/data source rather than inferred from the optical spectrum.

Each row of the regression dataset is therefore a sample-level spectral vector paired with a measured target or six-target vector. An extra wavelength is another feature, not another independent sample. With 136 observations and 1,921 inputs, the nominal feature-to-observation ratio is about 14.1, although neighbouring spectral bands are strongly correlated.

Reference values also carry assay uncertainty and depend on the method and reporting basis. A prediction should retain the target definition and units used in calibration. A detailed image of predicted composition would require additional spatially valid evidence, because these labels describe prepared samples rather than individual flour particles.

Sources: [1]

DLGX: organise the exploration before selecting a model

My confirmation presentation dated 28 April 2026 introduces DLGX as the Deep Learning Grid Explorer Framework for spectral-only data analysis. Its diagram identifies deep convolutional encoders, Vision Transformers and graph convolutional networks; configuration-driven experiments; configurable preprocessing; batch evaluation; consolidated metrics and visualisation; and single-target or vector prediction. This describes research exploration infrastructure used around the chickpea application, not a separate claim of a published software product.

The published experiment first compared two CNN encoder variants, a ViT and a GCN. CNN filters operate over local spectral neighbourhoods. The ViT partitions the spectrum into patches for attention. The GCN requires an explicit graph representation of the signal. These are different inductive biases and representations, not interchangeable model names.

The model-selection stage used an 80/20 train/test split with random state 42 and compared preprocessing choices under controlled settings. It evaluated separate models for each target as well as a model predicting all six targets together. The vector-prediction experiments were weaker in this dataset. Sharing an encoder across targets can help in some settings, but this study does not establish an advantage for joint prediction.

DLGX’s useful contribution to the workflow is making the comparison systematic: keep the sample identifiers, target definitions, preprocessing, architecture, hyperparameters and split attached to every result. A framework name does not itself guarantee leakage-free evaluation, and the confirmation presentation’s description should be distinguished from the paper’s published methods and outcomes.

Sources: [1]

What the chickpea experiments found

CNN Encoder 01 was selected for the more extensive evaluation against partial least squares regression (PLSR), using shuffled 10-fold cross-validation with random state 42. The paper reports four preprocessing pipelines, with PP-04 combining standard scaling, a translation and squaring. That empirical selection is specific to these data and experiments; it does not mean squaring spectra is a generally preferred NIR procedure.

The reported CNN comparison had higher average R² and lower average RMSE than PLSR across the six targets, with benefits that varied substantially by target. The paper’s rounded discussion gives soluble-sugar mean R² increasing from about 0.0 to 0.2 and mean RMSE decreasing from about 1.0 to 0.8 percentage points. For moisture, mean R² increased approximately from 0.7 to 0.8, while RMSE remained around 0.2 percentage points for both approaches. These rounded summaries are not exact fold statistics.

An R² around 0.2 remains modest even when it improves on a baseline. RMSE must be read in the target’s units and against the composition range, assay uncertainty and intended decision tolerance. The paper explicitly states that the achieved performance was not sufficient for industrial application as presented. These results support continued calibration research rather than an instrument-independent, deployment-ready flour analyser.

Sources: [1]

Validate the next calibration at the sample and batch level

For a follow-up evaluation, fit learned preprocessing inside each training fold and apply the fitted transformation to the held-out observations. Choose architectures and hyperparameters inside an inner validation loop, then estimate generalisation on an outer loop or a locked external test set. Repeated selection on the same folds can make the best reported configuration look stronger than its truly independent performance.

If there are repeated scans or repacked measurements from one flour sample, group them together in the same fold. For deployment, test new grain batches, growing environments, seasons, preparation conditions and instruments. Random sample folds from a common experimental population do not directly measure those forms of shift. These are recommendations for stronger future validation; they are not claims that the published work already performed nested or external evaluation.

Retain PLSR and simple baselines alongside neural models. Report per-target errors, fold variability and prediction-versus-reference plots, then inspect where the calibration fails. Finally, separate the three validations: does the instrument produce repeatable spectra, does the assay provide reliable reference values, and does the fitted model generalise to the samples on which it will be used?

Sources: [1], [3], [4]

Frequently asked questions

Is this flour measurement hyperspectral imaging?

No. The study measured one sample-level reflectance vector per observation. It did not acquire a spatially resolved image cube.

Why are there 1,921 features?

The reported grid covers 680 to 2,600 nm inclusive at 1 nm intervals: 2,600 − 680 + 1 = 1,921. That is the number of sampled wavelengths, not independent flour samples.

Is the 3D model the exact SpectraStar instrument?

No. It is a conceptual diffuse-reflectance measurement geometry. The paper supplies the instrument name and sampling range, but not enough internal optical or sample-cup details to reconstruct its hardware.

Does the animation predict flour composition?

No. Its curves and slider effects are synthetic. It contains no trained chickpea regression model and no real spectra from the study.

What was DLGX used for?

The confirmation presentation describes an exploratory spectral-learning framework with configurable models and preprocessing, batch comparisons, evaluation plots, and single or multiple regression targets. The chickpea study is its published application context.

Can I use the reported CNN on commercial flour immediately?

The paper states that performance was not sufficient for industrial application as presented. New populations, preparation procedures and instruments require representative calibration and independent validation.

References and further reading

  1. Zia, Husnain et al. (2025). Unlocking chickpea flour potential. Current Research in Food Science, 10, 101030.
  2. BÜCHI: NIR principles, overlapping absorption bands and calibration development
  3. scikit-learn: Cross-validation and grouped observations
  4. scikit-learn: Common pitfalls, preprocessing leakage and pipelines

Reader feedback

— reads— comments

Reads since 1 October 2026: at least 15 seconds with the article visible, counted once per browser per day. Reactions are anonymous and can be changed.

Discuss this article

Ask a technical question, challenge an assumption or share evidence from your own work.

Your name and comment will be public after approval. New comments wait for approval. No email required. Keep it respectful and relevant; no personal information or spam. Plain text, up to 2,000 characters.

Owner: manage comments