
A 2023 reading route organized around evidence
A useful geospatial representation has to survive more than a change of classifier. The input bands, the geography represented in pretraining, the downstream task and the tuning budget all shape the result. These notes use three NeurIPS 2023 papers to separate those decisions: GEO-Bench supplies an evaluation framework, CROMA supplies a multimodal objective, and SSL4EO-L supplies a sensor-specific data and pretraining resource. [1][2][3][4]
The scope is deliberately historical and selective: one main-conference paper, CROMA, and two Datasets and Benchmarks papers from volume 36. This is a reading guide checked on 2 October 2026, not a roundup of NeurIPS 2026, an attendance report or a ranking of every geospatial model. The selected full PDFs were read for methods, evaluation and stated limitations. Current repository caveats are separated from the original papers. [1]
Start with the decision you need to defend. Choosing an encoder requires one kind of evidence; deciding which sensor archive can support an application requires another. A reported improvement is useful only after its data access, adaptation procedure and measurement become clear. The reading order below begins with those contracts, so architecture names arrive after the experimental question.
GEO-Bench: read the evaluation rules before the score
GEO-Bench curates six classification and six segmentation tasks. The paper distinguishes modified datasets with an m- prefix, preserves original splits where available, and describes spatially non-overlapping splits otherwise. Its protocol recommends a bounded hyperparameter search, at least ten seeds for the selected configuration, normalized task scores and uncertainty-aware aggregation with an interquartile mean. Read §§3–4 before using a figure as a model-selection shortcut. [2]
The authors explicitly note missing coverage of temporal fine-tuning and fusion with text or weather, alongside incomplete geographic and biome coverage. Those limits matter when the desired application depends on a season, a weather record or a region underrepresented by the tasks. An aggregate improvement should therefore be followed by a task-level inspection. [2]
My first check would be a one-page evaluation contract: which benchmark version, which channels, frozen or fine-tuned encoder, which validation budget, and which per-task metrics? Keep raw seed results. A reader should be able to see whether a method wins broadly or whether a small number of tasks carry its average.
Sources: [2]
| Paper / track | Question and data | Method / protocol | Artifacts and first read |
|---|---|---|---|
| GEO-Bench / Datasets and Benchmarks | How should transfer be measured across curated EO tasks? | Task normalization, tuning budget and seed-aware aggregation | Official suite and CSVs; read §4, then current Known issues [2][5] |
| CROMA / Main | What transfers from aligned radar–optical data? | Contrastive agreement plus masked fusion; separate encoder uses | Official models; read pairing and modality contract in §2 [3][6] |
| SSL4EO-L / Datasets and Benchmarks | What does Landsat archive coverage permit? | Product-specific pretraining and segmentation evaluations | TorchGeo resources; read sampling §2.1 and limitations §5 [4][7] |
CROMA: separate shared information from fused information
CROMA combines radar–optical contrastive learning with masked reconstruction. Separate encoders process Sentinel-1 and Sentinel-2; a multimodal encoder fuses their representations. Its 2D-ALiBi and X-ALiBi biases encode spatial relationships in self-attention and cross-attention. The methods section makes the pairing assumption concrete: positive radar and optical samples are matched geographically and temporally. [3]
The paper evaluates several uses of the representation, including frozen-feature classifiers, fine-tuning and segmentation. Its stated limitation is a focus on static-in-time Sentinel-1/2 rather than other sensor families or time-series modeling. The official repository provides pretrained-model and encoder entry points; it was inspected as a resource, not executed here. [3][6]
For a new task, ask whether the useful signal is shared across the sensors or exists mainly in one of them. Then choose the corresponding branch and record which modalities are present at inference. An optical-only result, a radar-only result and a jointly fused result answer different deployment questions. They should remain distinguishable in the experiment log.
SSL4EO-L: treat the archive and sampling rule as a method
SSL4EO-L builds Landsat pretraining data across sensor/product combinations and evaluates learned representations through segmentation. The data construction samples around cities and asks for four seasonal observations, while excluding unsuitable patches. Read §2.1 alongside the downstream-data sections: historical coverage and preprocessing are central parts of the contribution. [4]
The paper states that the sampling rule underrepresents persistently cloudy and sparsely populated regions. It also notes that the new benchmark datasets are limited to the United States, which limits what their scores establish elsewhere. These are stated author limitations, not failures inferred from an abstract. [4]
TorchGeo exposes distinct product choices such as TM top-of-atmosphere and OLI surface reflectance. That interface is a useful reminder to select a checkpoint with a matching measurement contract. Before adapting one, list the band order, product level, scaling and season handling. A tensor with the expected dimensions can still carry the wrong physical meaning. [7]
A repository check changes the reproduction plan
The current GEO-Bench repository documents incorrect band-name and wavelength metadata in classification_v1.0 for m-eurosat and m-brick-kiln. It says the pixel data are unchanged, but selecting channels by the erroneous names can select the wrong channel. For m-brick-kiln, the final three channels are true-colour composites rather than reflectance bands. The repository supplies the actual ordering and warns that hosted data have not simply been regenerated with the correction. [5]
This changes the first action: inspect the version and channel mapping before training. Do not repair a metadata issue by reordering values according to the incorrect names. Save both the on-disk order and the semantic order supplied to the model, with the mapping source. If comparing an old result with a corrected loader, disclose that protocol difference rather than attributing every change to the model.
The same repository directs new users toward a successor benchmark. These notes retain the 2023 paper as the reading subject; they do not silently substitute a newer task suite. Pin the intended benchmark, loader and checkpoints. Reproduction of an old table and evaluation on a current benchmark can both be worthwhile, but they need separate labels. [5]
Sources: [5]
What can actually be compared?
The comparison table maps the role of each paper rather than merging their reported numbers. A dataset collection, an encoder objective and an evaluation framework occupy different positions in an experiment. Reading them together is useful because it makes missing decisions visible: where observations come from, how features are learned, and how transfer is measured.
For an independent comparison, first fix the target question. Is the goal label efficiency in one region, robust transfer across regions, or performance under a missing sensor? Keep the target split and metric stable while changing one major design choice. Record any extra pretraining corpus, modality, label access or tuning that a candidate receives. Otherwise an architecture comparison can accidentally become a data-budget comparison.
Use both a task-level view and a summary. Inspect rare classes and failure groups relevant to the application before deciding that an average is adequate. If two models trade places across tasks, present the trade-off. An overall winner is a stronger claim than a useful candidate for a particular workflow.
For an HSI project, this selection offers protocol ideas rather than direct proof of narrow-band material discrimination. None of these selected papers establishes a universal hyperspectral sensor interface. Translate a lesson such as versioned band handling or geographic holdout into the HSI measurement setting, then state exactly which transfer claim remains untested.
A compact reading session with a concrete output
Read GEO-Bench §4 first and draft the evaluation contract. Next, read CROMA’s objective construction and note which data relationships create positives, targets and fused tokens. Then read SSL4EO-L’s sampling and limitations sections and mark which places, seasons and products your application needs. Finally, revisit the source repositories for version-specific issues. This ordering is editorial guidance, not a prescribed author workflow.
End the session with an evidence matrix, not a list of impressive scores. One row should contain the hypothesis, admissible data, checkpoint identity, adaptation budget, held-out group, metric, and a result that would contradict the hypothesis. Add an access column: source inspected, data download checked, environment installed, or experiment reproduced. Only the first of those stages is completed by these notes.
A sensible stopping point is a fully specified small comparison ready for review. Large downloads and training can wait until the protocol is coherent. The payoff from this selection is knowing which evidence would make a model useful for the intended task, and which attractive claims the current reading still cannot support.
Download this reading guide and original cover (ZIP)
Frequently asked questions
Which edition and tracks are covered?
Three NeurIPS 2023 papers: CROMA in the main conference; GEO-Bench and SSL4EO-L in the Datasets and Benchmarks track. This is a selected reading route, not comprehensive coverage.
Were the full papers actually read?
Yes. The official PDFs were inspected for the selected methods, protocol and limitations sections. Code and dataset listings were checked where stated; no training result was reproduced.
Can the three papers share one leaderboard?
No common experiment is established here. One contributes an evaluation suite, one a multimodal learning method, and one Landsat data and pretraining resources.
Does CROMA establish HSI transfer?
Its selected paper studies Sentinel-1/2 representations. Applying its objective or spatial biases to hyperspectral data requires a separate sensor-specific evaluation.
Why check GEO-Bench channel metadata?
The current official repository documents band-metadata problems in two classification_v1.0 tasks. Its mapping and warnings should be checked before selecting channels by name.
Should a new run use the 2023 code automatically?
Choose intentionally. Pin historical artifacts when reproducing the 2023 protocol, or label newer loaders and benchmark versions as a new evaluation.
References and further reading
- NeurIPS 2023: official proceedings, volume 36
- Lacoste et al. GEO-Bench: Toward Foundation Models for Earth Monitoring. NeurIPS 2023, Datasets and Benchmarks; §§3–4, §7.
- Fuller, Millard and Green. CROMA. NeurIPS 2023, main conference; §§2–3, §6 and appendix.
- Stewart et al. SSL4EO-L: Datasets and Foundation Models for Landsat Imagery. NeurIPS 2023, Datasets and Benchmarks; §§2–5.
- GEO-Bench official repository: versioned loading instructions and Known issues
- CROMA official repository: encoders and pretrained models
- TorchGeo: SSL4EO-L dataset documentation
Reader feedback
Reads since 1 October 2026: at least 15 seconds with the article visible, counted once per browser per day. Reactions are anonymous and can be changed.
Loading reader feedback…
Discuss this article
Ask a technical question, challenge an assumption or share evidence from your own work.