
Turn one claim into a question you can actually test
A conference paper can suggest a useful experiment without supplying a complete experimental recipe. Before installing anything, choose one claim, locate its supporting evidence, and write the conditions under which your result would count as consistent, inconsistent or unresolved. A successful command is a software event. A supported claim needs a defined comparison.
This guide connects the existing conference reading routes to a small executable experiment contract. The running example asks whether a fixed, finite-context image operator gives equivalent outputs when evaluated eagerly or in tiles. It reuses the exact NumPy operator from the out-of-core HSI guide, adding a frozen card, negative controls, decisions and an inspectable run receipt.
The experiment is synthetic and CPU-only. It uses no trained network, observed hyperspectral scene or ground-truth labels. Its connection to a published conference paper is explicitly limited to an overlap-tile idea. The download does not reproduce a paper’s architecture, training, accuracy or speed.
Say what reproduction, replication and a toy test mean here
Terminology differs across research communities. This guide follows the National Academies’ convention: computational reproduction seeks consistent results with the same data, methods and computational conditions; replication addresses the same scientific question with independently obtained data. State your chosen convention rather than assuming readers use those words identically. [1]
A reimplementation is a separate engineering description: you wrote a new implementation of a method. It can contribute to a reproduction or replication, but that depends on which data, protocol and claim it tests. Porting an idea into a different architecture does not automatically establish either outcome.
A toy mechanism test deliberately reduces the problem until a particular assumption can be inspected. Here it checks local context, boundary handling, preprocessing consistency and output coordinates. A passing result supports the stated synthetic comparison. It does not demonstrate that a published model reproduces its reported result, that a real material can be classified, or that the approach transfers to another sensor.
Sources: [1]
| Study | What is held or changed | Evidence needed | What it can support |
|---|---|---|---|
| Computational reproduction | Original data and computational recipe, with deviations disclosed | Pinned artifacts, executable conditions, original metric and consistency rule | Consistency of the specified original calculation [1] |
| Replication | Same scientific question, independently obtained data | Sampling design, uncertainty and justified comparison | Consistency across independent studies [1] |
| Reimplementation | New code for a described method | Method correspondence, deviations and numerical tests | Implementation evidence; study type still depends on data and protocol |
| This toy mechanism test | Original small operator; synthetic data; controlled execution changes | Saved arrays, declared tolerances, positive and negative controls | Fixture-specific finite-context equivalence only |
Trace the motivating paper to a specific passage
Ronneberger, Fischer and Brox’s U-Net paper appeared at MICCAI 2015. The inspected text is the explicitly versioned arXiv manuscript 1505.04597v1, dated 18 May 2015; the publisher’s chapter is separately identified by DOI. Page 3 and Figure 2 explain overlap-tile segmentation with sufficient input context and mirrored missing border input. Page 4 also constrains tile geometry for pooling. Those passages motivate this example. [2][3]
Write the claim location into the card, along with authors, venue, manuscript version and a stable source link. Record which supplement or later correction you actually used. A method description, an ablation and a leaderboard entry are different evidence objects; each may have a different date or implementation.
The bridge made here is an explicit reduction: replace the published network with a transparent radius-four operator, then test whether its declared context is enough. Its available-neighbour boundary rule and final decimation are part of the teaching contract. This reduction makes a useful software property easy to check while leaving the original paper’s empirical claims untested.
Runnable experiment card · synthetic CPU evidence
Inspect the claim-to-test contract
Six actual calculations check one fixed finite-context operator. Two sufficient-halo cases match; four negative controls differ as planned. No published model, dataset or performance result is reproduced.
Default input: 67 × 83 × 16 float32. Output: 34 × 42 × 3. Core: 16 pixels. Global stride: 2. Tolerances: atol 0.000002 and rtol 0.00001, relative to the eager reference.
| Case | Maximum logit error | Label disagreement | Decision |
|---|---|---|---|
| Halo 0, aligned | 0.203161 | 3.431% | Different, expected |
| Halo 2, aligned | 0.098226 | 0.630% | Different, expected |
| Halo 4, aligned | 0.000000 | 0.000% | Equivalent, expected |
| Halo 7, aligned | 0.000000 | 0.000% | Equivalent, expected |
| Halo 5, wrong phase | 0.202371 | 6.863% | Different, expected |
| Halo 4, per-tile centering | 0.844525 | 27.941% | Different, expected |
Zero recorded error occurred on the tested environment. The decision uses the stated tolerance; it does not promise zero error on every platform.
1. Read the scope before running
U-Net’s overlap-tile discussion motivates the question. This test reuses a much smaller fixed NumPy operator from the portfolio’s tiling guide. It has a four-pixel support radius, its own per-layer available-neighbour edge rule and no learned parameters. Supervised splits are not applicable. A full U-Net reproduction remains an unaudited separate project.
2. Run the card in a fresh directory
Extract the ZIP and follow README.md. Tested with CPython 3.12.14 and NumPy 2.3.5 on Linux x86_64. Install the pinned dependency if necessary; all experiment steps then run offline.
python python/experiment.py validate --card experiment-card.json
python -m unittest discover -s tests -v
python python/experiment.py run --card experiment-card.json --out my-run
python python/experiment.py verify --out my-runChoose a new output path. Existing directories are refused. Exit 0 means all case expectations matched; exit 1 preserves an expectation mismatch; exit 2 reports a handled input, output or verification error.
3. Understand a negative outcome
A planned difference is a useful control. An unexpected agreement or disagreement needs investigation. Keep the original card, arrays and report. Diagnose coordinates, edges, preprocessing and arithmetic before revising the hypothesis or tolerance. A modified card receives a different digest; it does not overwrite the earlier receipt.
4. Check the limits of the receipt
The verifier checks the file set and hashes, then recalculates metrics from saved arrays. It does not authenticate the publisher, prove preregistration, reproduce the paper, or validate software and data rights. Read-window bytes exclude intermediates and process memory; local timings are not a speed benchmark.
Canonical contract SHA-256: 6f25424ac8fc51976dabffb15c4f44dc258c4b964b866c1d57e17dae326b8d07
For the surrounding workflows, use the tile-and-halo guide, run ledger and pipeline capstone.
Audit code, data and rights before committing to a reproduction
The authors’ official project page names u-net-release-2015-10-02.tar.gz and describes a Caffe-based release with trained networks and a MATLAB overlap-tile interface, tested on Ubuntu 14.04 and MATLAB 2014b. That is a concrete historical artifact lead. The archive was not downloaded or executed for this guide, so its file digest, complete component licences and current executability remain unverified. [4]
The selected manuscript uses arXiv’s non-exclusive distribution licence. That grant to arXiv should not be treated as a blanket reuse licence for paper figures, software, weights or datasets. The present bundle links the paper and includes no third-party PDF or figure. The original benchmark data and their licences were not audited, and the card therefore keeps the full-reproduction gate unresolved. [5]
For a real attempt, keep separate records for source code, dependencies, pretrained weights, images, labels and metadata. Give each an exact version or immutable identifier, retrieval location, checksum where obtained, licence text or verified licence reference, access conditions and any redistribution restriction. Mark unknown fields as unknown. A downloadable file and a repository licence do not establish the rights for every linked resource.
The executable in this download is original portfolio teaching code; its core operator is preserved byte-for-byte from the earlier tiling article. Generated arrays are synthetic project artifacts. No additional reuse licence has been selected for those original assets; RIGHTS.md states that plainly. NumPy is a separately installed dependency under its documented BSD-style licence, and is not vendored into the ZIP. [6]
Write the contract before looking at the outcome
A useful contract connects the claim, intervention, comparison, evidence and decision rule. It should let someone distinguish a planned test from a convenient explanation invented after the run. The supplied JSON card has a schema version, experiment ID, source provenance, input specification, operator identity, case matrix, metric rules and compute bounds. Its canonical digest travels into the report.
The baseline is eager evaluation of the full generated cube. The intervention changes the execution partition, halo, output-grid alignment or preprocessing scope. Every case compares all output logits against that one reference. Two sufficient-halo cases should agree; four deliberately altered cases should differ. A successful demonstration requires every declared expectation and every exactly-once output-write check to hold.
The default input is 67 × 83 × 16, in row–column–band order, with float32 values. The output is 34 × 42 × 3 because stride-two sampling starts at global coordinate zero. These shapes, a 16-pixel core, the tolerances and six planned cases are readable in the card. Editing the card creates a different contract digest; it does not quietly update the previous run.
This is an executable specification, not independently timestamped preregistration. A file can be edited before publication, and a hash does not prove when its contents were chosen. If prospective registration matters, use an appropriate independent archive or registration process and preserve the original plan plus dated amendments.
Give data, splits and preprocessing separate entries
In the toy, values come from a coordinate-defined generator with no random-number stream. The same coordinate returns the same spectrum regardless of the window that requested it. There are no physical wavelengths, calibrated units or real land-cover classes. The three output channels are fixed teaching logits; taking their argmax creates arbitrary labels rather than reference truth.
No parameter is fitted, so train, validation and test splits are explicitly not applicable. Creating artificial split names would add scientific theatre without a learning boundary. The cases test one deterministic calculation on one synthetic fixture. Multiple tile sizes and additional geometry tests expand software coverage, but do not create independent scenes or statistical replications.
For an actual HSI learning study, replace that entry with scene/acquisition identity, sensor, units, wavelength order, label ontology, invalid-pixel handling and saved membership masks. State whether splitting occurs by pixel, patch, field, flight or scene. Account for the full support of spatial preprocessing and model inputs, then decide whether the claim requires independence at an even larger acquisition scale.
Fit learned scaling, dimensionality reduction and feature selection only inside the authorised training partition, unless the declared task intentionally permits another access policy. Record that exception rather than hiding it. The spatial-split guide and pipeline capstone already provide the detailed split and fit-boundary workflow; this contract links to those checks instead of rebuilding them.
Derive the finite-context claim and its limits
The reused operator applies fixed per-band standardisation, a projection to four channels, two 5 × 5 available-neighbour means with tanh nonlinearities, and a projection to three logits. A final step samples every second pixel. Each mean adds a radius of two, so a dense output depends on inputs at most four pixels away in each spatial direction. Pointwise transforms and projections add no spatial support.
An aligned tile can therefore retain a core whose required in-domain inputs are present in its read window. Internal tile-edge effects are discarded with the halo; real image edges retain the same per-layer rule as the reference. Output ownership and global sampling phase must also agree. These conditions explain the radius-four sufficiency claim for this particular operator.
The code’s left/top alignment may expand a read beyond its nominal halo. A smaller halo can occasionally cover all inputs needed for particular sampled outputs. The guide claims that four is sufficient here, not that every smaller setting fails for every geometry, or that four is a universal minimum.
Global attention, scene-wide normalization, recurrent state, resampling and encoder–decoder paths require their own dependency and coordinate analysis. A locally computed mean changes the operator when it replaces fixed preprocessing. Giving that changed operator a larger halo cannot establish equivalence to the original calculation merely because the stitched image looks smoother.
Choose metrics that can falsify the narrow claim
The primary rule checks every finite logit: absolute(actual − reference) must be at most 2 × 10⁻⁶ + 10⁻⁵ × absolute(reference). The comparison passes the reference as NumPy’s second argument and checks equal shapes first, avoiding accidental broadcast comparisons. NumPy documents both that asymmetry and why a default absolute tolerance can be unsuitable near zero. The tolerances here are explicit regression tolerances for this float32 teaching calculation. [7]
The report also retains maximum absolute error, mean absolute error and argmax-label disagreement. The first makes a localized defect hard to average away. The second describes typical numerical deviation. Label disagreement shows whether an implementation difference changes the selected output channel. It is not classification error because no ground-truth labels exist.
Close logits can still swap an almost-tied argmax, so equivalent logits do not logically guarantee identical labels. Conversely, matching labels can hide large logit differences. The tests contain a near-tie example to preserve that distinction. Always inspect the primary rule and the secondary diagnostic under their own meanings.
For a paper benchmark, retain the paper’s metric definition, units, class mapping, averaging, ignore-mask convention and official evaluator version. A new metric can be useful as an additional analysis, but replacing the reported metric prevents a direct numerical comparison. Define an uncertainty and consistency criterion suitable for the actual sampling units and seed design.
Sources: [7]
Keep the negative controls visible
The executed default fixture gives maximum logit error 0 with halos four and seven. Halo zero differs by approximately 0.20316, halo two by 0.09823, the deliberately wrong output phase by 0.20237, and per-tile centering by 0.84453. These are outputs of the supplied synthetic calculation, not copied or invented paper scores. The companion table lists the associated label disagreements.
A case expected to differ is a successful negative control when a difference is observed. It should remain labelled “different, expected” rather than a failed research run. If a sufficient-halo case differs, or a negative control unexpectedly agrees, the overall outcome becomes expectation-mismatch; the runner retains the result and returns a nonzero status.
That mismatch needs diagnosis before scientific interpretation. Check shapes, global coordinates, preprocessing scope and boundary rules first. Then inspect the arithmetic and tolerance rationale. Do not silently loosen the tolerance until a desired result passes. Preserve the original card and report, explain the defect or revised hypothesis, and run the amended contract into a new directory.
A shared bug in eager and tiled evaluation could still produce agreement. The suite therefore also includes a direct small-array neighbourhood-mean oracle, a distant-input perturbation check and tests for edges, odd shapes, strides, finite values and exactly-once writes. These tests improve coverage while remaining short of a formal proof for arbitrary programs.
Plan the compute and stop conditions before scaling
The default source needs a local CPU, CPython and NumPy. It performs one eager evaluation and six bounded tiled cases, with no optimization, GPU allocation, remote service or dataset download. The main input contains 355,904 bytes. The parser caps spatial dimensions, bands, halo, core and number of cases; it also enforces the card’s input-byte ceiling before allocating the fixture.
The report records per-case elapsed time and maximum read-window bytes. Those values have deliberately modest meanings: timings are single-run diagnostics, and read bytes exclude intermediates, Python, libraries, output storage and process RSS. This small run does not establish out-of-core scalability, accelerator savings or a speedup over a published implementation.
For a larger reproduction, write a staged budget: artifact inspection, one forward pass, tiny overfit check where training applies, one complete baseline, then the predeclared method and seed matrix. Include expected memory, wall time, storage, access costs and failed-run allowance. Set a stop condition for an unresolved artifact, violated split contract, numerical fault or exhausted budget before launching the expensive stage.
The NeurIPS checklist explicitly asks researchers to describe experimental settings, uncertainty and compute, including resources spent beyond the final reported runs. It provides a useful reporting prompt here; this article makes no checklist-compliance or venue-acceptance claim. For stochastic frameworks, record more than seeds: PyTorch also warns that reproducibility is not guaranteed across releases and platforms. Neither PyTorch nor a GPU was used for this experiment. [8][9]
Run the experiment card and inspect the receipt
Download the complete runnable source ZIP, extract it, and follow README.md. The tested environment is CPython 3.12.14 and NumPy 2.3.5 on Linux x86_64. requirements.txt pins the tested NumPy release; this is not a complete portable environment lock. The program runs offline once that dependency is available.
From the extracted folder, use python python/experiment.py validate --card experiment-card.json, then python -m unittest discover -s tests -v. Run python python/experiment.py run --card experiment-card.json --out my-run, followed by python python/experiment.py verify --out my-run. The README includes optional single-thread environment settings used for the recorded example. Always choose a new output directory.
The completed run has a card, generated input, eager reference, six output arrays, six write-count arrays, environment record, report and exact executed source snapshots. A final manifest checksums 19 files. Verification checks the file set and digests, then recomputes the reported metrics from saved arrays. It does not execute a snapshot as code.
An unsigned manifest establishes consistency with those recorded bytes, not publisher authenticity or prospective registration. A crash can leave a partial directory without a completion manifest; retain it as an incomplete attempt and use a new directory when retrying. The experiment-tracking guide explains how to keep attempts, retries, selection decisions and final tests connected.
Finish with an honest conclusion and a separate next study
The supported conclusion is narrow: on the declared synthetic fixture and tested environment, the correctly aligned fixed-preprocessing cases with sufficient halo matched the eager logits under the declared tolerance, while the selected negative controls differed. The code and receipt let another reader rerun that comparison and investigate disagreement.
A full paper reproduction remains a separate project. It needs the audited original artifacts or an explicitly documented reimplementation, an executable historical or justified replacement environment, authorised datasets and weights, the relevant split and preprocessing recipe, and the correct evaluation procedure. Report any deviation before comparing numbers. Missing access is a blocked attempt, not evidence that the scientific claim is false.
To choose that next project, use the existing IGARSS notes, NeurIPS notes, ICCV notes and workshop notes. Select one contribution whose artifacts and compute fit the question. Carry forward the same discipline: a precise claim, an explicit contract, an informative negative outcome and a conclusion no broader than the evidence.
Frequently asked questions
Does this reproduce U-Net or a published HSI result?
No. U-Net motivates the overlap-tile question. The executable uses the existing synthetic finite-context NumPy operator, with different geometry and edge handling, no fitted weights and no observed data.
Why are there no train, validation and test splits?
Nothing is learned or selected from labels in this deterministic mechanism test. Those splits are explicitly not applicable. A real learning study needs its own saved memberships, fit boundary and selection protocol.
Do the expected differences count as failures?
A negative control is successful when its predeclared difference appears. An unexpected agreement or disagreement creates an expectation-mismatch outcome, retained in the report with a nonzero run exit status.
Does a four-pixel halo work for every model?
No. Four follows from the two radius-two local operations in this fixed operator. Other architectures, padding, global operations, preprocessing or output geometry need a new dependency analysis.
Can I rerun the source without a GPU or account?
Yes. The source needs Python and the pinned NumPy dependency, then runs offline on the CPU. It downloads no model or dataset and refuses to overwrite an existing run directory.
What do the hashes and tests establish?
They check source identity, artifact consistency, saved-output metrics and the tested numerical properties. They do not authenticate the publisher, prove preregistration, establish licence rights or validate a broader scientific claim.
References and further reading
- National Academies of Sciences, Engineering, and Medicine (2019). Reproducibility and Replicability in Science, Summary: definitions. DOI 10.17226/25303.
- Ronneberger, Fischer and Brox (2015). U-Net: Convolutional Networks for Biomedical Image Segmentation. Inspected manuscript arXiv:1505.04597v1, pp. 3–4.
- Ronneberger, Fischer and Brox (2015). U-Net, MICCAI 2015, LNCS 9351, pp. 234–241. Publisher record; DOI 10.1007/978-3-319-24574-4_28.
- University of Freiburg, LMB. Original U-Net project and u-net-release-2015-10-02 archive description.
- arXiv. Non-exclusive licence to distribute, version 1.0.
- NumPy Developers. NumPy licence, v2.3 documentation.
- NumPy Developers. numpy.allclose, v2.3 documentation: reference-relative tolerance and broadcasting.
- NeurIPS. Paper Checklist Guidelines: experimental settings, uncertainty and compute reporting.
- PyTorch Contributors. Reproducibility: randomness and release/platform limitations.
Reader feedback
Reads since 1 October 2026: at least 15 seconds with the article visible, counted once per browser per day. Reactions are anonymous and can be changed.
Loading reader feedback…
Discuss this article
Ask a technical question, challenge an assumption or share evidence from your own work.