Original three-panel diagram of nested spatial supports, diverse supervision layers and one positive point among unlabeled points, illustrating three representation-learning contracts.
Original editorial schematic. Shapes are illustrative, not measured data, published figures or model results.

Representation learning has three different bottlenecks

What prevents a remote-sensing model from being useful: uncertainty about physical scale, limited breadth of supervision, or a shortage of trustworthy negatives? This reading route assigns one ICCV 2023 paper to each question. Scale-MAE addresses scale-aware pretraining, SatlasPretrain broadens the types of remote-sensing labels, and T-HOneCls studies positive-unlabeled hyperspectral learning. [2][3][4]

All three are from the ICCV 2023 main proceedings, not the workshops. The official CVF PDFs were inspected for selected method, evaluation and discussion sections. The guide was checked on 2 October 2026; its edition is intentionally fixed, and it makes no attendance or reproduction claim. This is not an exhaustive Earth-observation survey or a claim about what won a later conference. [1]

The common thread is the contract between an observation and its learning signal. A pixel spacing tells a model something different from a category label; a known positive tells it something different from an observed negative. Reading these papers together helps prevent those distinctions from disappearing inside the broad phrase “better representation.”

Sources: [1], [2], [3], [4]

ProceedingsScopeRead papersCompareReading list
Conceptual workflow. Each stage requires its own assumptions and checks.

Scale-MAE: physical support belongs beside pixel dimensions

Scale-MAE modifies masked-autoencoder pretraining with ground-sample-distance-aware positional encoding and a decoder targeting low- and high-frequency image components. The goal is to learn representations across known spatial scales. Read §3’s construction before inspecting the transfer tables: the absolute scale information is part of the input, rather than an incidental image resize. [2]

The discussion notes an important limit: the model stacks input bands and does not directly handle bands that all have different ground sample distances. It also reports higher GPU memory use at an equal batch size despite a smaller decoder. Those details keep “scale-aware” from becoming a claim of universal multisensor compatibility or automatically cheaper training. [2]

My first check would trace GSD through every crop and resize in a downstream pipeline. Retaining an old metadata value after resampling can describe the wrong spatial support. A useful follow-on study would distinguish synthetic downsampling from a genuinely different sensor acquisition; agreement in the first setting would not settle the second.

Sources: [2]

ICCV 2023 methods solve different bottlenecks. Read-first questions are editorial; the table does not compare numerical performance.
PaperQuestion / dataMethod and evidenceFirst read / artifact
Scale-MAEHow does physical scale enter geospatial features?GSD-aware MAE; multiscale classification/segmentation transfer§3 input/decoder and §5 limits; repository is a reimplementation [2][5]
SatlasPretrainWhich structured supervision transfers?NAIP and Sentinel-2 modes; shared features and seven head types§3 label provenance, §4 heads; project linked, downloads not validated [3][6]
T-HOneClsHow can positives plus unlabeled HSI train a binary model?Taylor variational loss; EMA/KL teacher; limited-positive protocol§3 objective and §4 sampling; code inspected, external data not downloaded [4][7]

SatlasPretrain: supervision is a structured resource

SatlasPretrain organizes imagery and labels by geographic tile and time. Its seven label types include segmentation, regression, object geometry, properties and classification. SatlasNet uses a shared Swin backbone, temporal pooling and task-specific heads. The NAIP high-resolution and Sentinel-2 low-resolution image modes are trained and evaluated separately. Read §§3–4 to understand those modes and the label structure. [3]

The data discussion describes missing-label concerns in OpenStreetMap-derived categories and additional test-set correction for selected categories. The paper also distinguishes slow-changing labels from dynamic ones. Those details are as important for a reproduction as the headline dataset size: a missing object and a true negative are different observations. [3]

My read-first question is which label type gives the desired transfer signal. A segmentation head, an object detector and a property predictor make different demands on the shared features. For a fine-grained material task, write down the correspondence between the pretraining target and the downstream distinction instead of assuming that more categories automatically supplies the needed discrimination.

Sources: [3]

T-HOneCls: unlabeled does not mean negative

Zhao and colleagues study a binary positive-unlabeled HSI setting without requiring a supplied class prior. T-HOneCls uses a finite Taylor expansion of a variational objective to moderate the unlabeled-data gradient contribution, plus an exponential-moving-average teacher and a KL consistency term. Read §§3.1–3.2 as two linked interventions: the loss changes optimization pressure, while the teacher stabilizes training. [4]

The paper evaluates five HSI datasets and two RGB datasets; the HSI experiments use limited positive examples. Its order ablation shows why the expansion order and training dynamics matter. This is not ordinary fully labeled multiclass classification, and its results should not be moved into that category without changing the problem definition. [4]

For a new target, my first question would be how positives were selected. Easy, well-lit or central examples may not represent the positive population encountered later. The absence of a user-specified prior does not eliminate this selection question. Predeclare how a decision threshold and stopping rule will be chosen without repeatedly inspecting the final evaluation labels.

Sources: [4]

Read the three interventions in a useful order

Start by writing the output you need: a scene category, an object boundary, a property value or a binary material mask. Next list the available supervision: complete labels, partial labels, positive-only annotations or unlabeled imagery. Then record physical support and acquisition metadata. This short inventory makes the first paper choice much easier.

If the same object occupies very different pixel extents, begin with Scale-MAE’s input construction. If a shared backbone needs varied tasks, begin with the Satlas label taxonomy and prediction heads. If confirmed absences are hard to obtain, begin with T-HOneCls’s learning assumptions. These choices are editorial recommendations about reading order; they do not assert that any method will win the intended application.

Only then compare the evaluation. Identify what is frozen, what is tuned, which annotations are available, and what the metric rewards. A pretraining paper’s transfer table and a positive-unlabeled detection table can both be persuasive while supporting different claims. The comparison below records those distinctions instead of inventing a shared leaderboard.

Reproduction starts with the released artifact, not its name

The Scale-MAE repository explicitly describes itself as a reimplementation of code originally optimized for the authors’ distributed cluster. Its README supplies a checkpoint and evaluation routes, including alternative split-file behavior. Pin the code and choose explicit split files before treating a local run as equivalent to the paper. The repository was read here; its environment and checkpoint were not executed. [5]

The Satlas paper links a project site for data, code and weights. That site was reachable during this check but exposed too little machine-readable content to verify every artifact download. The PDF is the evidence for the method notes; current end-to-end download and environment readiness remain unchecked. [3][6]

The T-HOneCls repository contains configurations, training scripts and requirements. It points to an external data service, which was not downloaded or authenticated to here. A visible repository establishes where to start inspecting the implementation; it does not establish that every dataset, dependency and result is immediately reproducible. [7]

For each artifact, record four separate states: paper inspected, source inspected, environment validated, and result reproduced. Avoid collapsing them into “available.” This makes time estimates more honest and gives a collaborator a clear place to resume. It also prevents an installation problem from being misreported as a scientific contradiction.

Sources: [3], [5], [6], [7]

A better next experiment isolates one uncertainty

Consider a hypothetical rare-material mapping task with a small set of confirmed examples and imagery collected at more than one resolution. It would be premature to combine every idea in this guide. First determine whether the bottleneck is missing supervision, unstable spatial support or insufficient features. A combined system may eventually be useful, but an initial study should make its primary intervention interpretable.

For a scale study, hold the annotation budget and split fixed while checking the metadata and input-resampling path. For a supervision study, hold the input representation fixed while defining exactly which labels are exposed. For a transfer study, preserve the downstream protocol and document the pretraining corpus. Choose one of these contracts as the central claim, and record the others as controls.

Define a counterexample before running anything. A scale-aware model might fail on an unseen acquisition; broad pretraining might miss the rare distinction; a positive-unlabeled model might become sensitive to how positives were sampled. These are possible failure modes to investigate, not findings attributed to the selected authors.

The useful outcome of the reading session is a bounded comparison with a meaningful negative result. If no such result could change the conclusion, the plan is still too vague. Keep the paper’s demonstrated setting, the released artifact and the proposed extension on separate lines. That habit turns an appealing conference contribution into evidence that can be checked.

Download this reading guide and original cover (ZIP)

Frequently asked questions

Are these ICCV main-conference papers?

Yes. All three selected PDFs are official CVF open-access versions of ICCV 2023 main-proceedings papers.

Is Scale-MAE a universal spectral adapter?

The selected paper addresses physical spatial scale. Its discussion explicitly notes restrictions around bands with different GSDs; that is not a general hyperspectral sensor interface.

Does SatlasPretrain mean self-supervised learning?

Its central resource contains diverse labels, and SatlasNet learns task-specific outputs. Describe the actual supervision used instead of labeling all pretraining self-supervised.

Is positive-unlabeled learning ordinary classification?

It has a distinct annotation contract: confirmed positives and unlabeled data are available, without assuming all unlabeled examples are negative.

Were the code packages executed?

No. The official PDFs and the stated resource documentation were read. Installation, full dataset retrieval and training remain separate unverified steps.

What should be read first?

Choose by the uncertainty: Scale-MAE for physical scale, SatlasPretrain for label/task breadth, and T-HOneCls for positive-only supervision. Then inspect the matching evaluation protocol.

References and further reading

  1. CVF: ICCV 2023 main-conference open-access proceedings
  2. Reed et al. Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning. ICCV 2023, §§3–5.
  3. Bastani et al. SatlasPretrain: A Large-Scale Dataset for Remote Sensing Image Understanding. ICCV 2023, §§3–5.
  4. Zhao et al. Class Prior-Free Positive-Unlabeled Learning with Taylor Variational Loss for Hyperspectral Remote Sensing Imagery. ICCV 2023, §§3–4.
  5. BAIR Climate Initiative: Scale-MAE official reimplementation and checkpoint documentation
  6. Allen Institute for AI: SatlasPretrain project
  7. T-HOneCls official implementation: configurations and requirements

Reader feedback

— reads— comments

Reads since 1 October 2026: at least 15 seconds with the article visible, counted once per browser per day. Reactions are anonymous and can be changed.

Discuss this article

Ask a technical question, challenge an assumption or share evidence from your own work.

Your name and comment will be public after approval. New comments wait for approval. No email required. Keep it respectful and relevant; no personal information or spam. Plain text, up to 2,000 characters.

Owner: manage comments