Skip to content

geobench 1.1.0

Latest

Choose a tag to compare

@nilsleh nilsleh released this 28 Sep 14:11
· 1 commit to main since this release
0051f44

Security

Loading a .tif, .hdf5, .npz or task_specs.pkl file no longer runs arbitrary code. Band metadata and task specifications are stored as pickles, and 1.0.0 read them back with plain pickle.loads, so a file from a third party could execute code on load (GHSA-jhvm-59f5-r2h6, GHSA-9rqp-gv68-cchq, GHSA-47r9-f95h-9xq3, reported by @adamjstewart). Unpickling now accepts only the types geobench writes (#32). The on-disk format is unchanged, so the published v1.0 data loads as before and does not need to be downloaded again.

Installation

  • Python 3.12 or newer is required, and 3.12, 3.13 and 3.14 are supported. 1.0.0 capped Python below 3.13.
  • The upper bounds on numpy, pandas, scipy, rasterio and the other dependencies are gone. 1.1.0 was tested against numpy 2, pandas 3 and torch 2.
  • matplotlib, pandas and seaborn moved to the plot extra, and torch and pytorch-lightning to the torch extra. Install with pip install "geobench[plot,torch]" if you use geobench.plot_tools or geobench.torch_toolbox.
  • affine is capped below 3, because affine 3 silently rebuilds the pickled v1.0 transforms with no coefficients set (#40).

Fixes

  • geobench-download no longer overwrites $GEO_BENCH_DIR in the process environment, and running geobench_download.py directly no longer tries to write to /mnt/home (#36).
  • Importing geobench no longer creates $GEO_BENCH_DIR, so the import works where that path is not writable (#36).
  • task_iterator(benchmark_dir=...) now loads datasets from benchmark_dir instead of from $GEO_BENCH_DIR (#38, fixes #22).
  • Writing a segmentation sample no longer appends the label to the caller's sample.bands. Writing the same sample twice used to raise on hdf5 and duplicate the label band on npz (#37).
  • GeobenchDataset instances are no longer kept alive by a method-level lru_cache (#37).
  • Smaller fixes in dataset.py: date formatting, the Landsat8 cirrus band name, masked-array input, and the np.float_ annotation removed in NumPy 2 (#37).
  • m-eurosat and m-brick-kiln store the wrong band names for most channels. Loading them now emits a warning, and the README lists the correct channel order (#30).

Behaviour changes

  • A missing benchmark directory now raises FileNotFoundError from task_iterator instead of being created empty on import.
  • ignore_task now drops the tasks it names; before, it kept only those tasks.