fix(data_catalog): reading from private bucket - #1536
Conversation
…ccess - Add handling for private bucket access in readers.py and xarray/zarr drivers. - Add aws profile to tests.yml - Add gh secrets for the Access/Secret keys. - Add boto3 and dask[distributed] to deps. - Added `open_zarrs` & `open_mfdataset` to `hydromt.readers`.
…terio driver for when caching s3 data locally.
…and call that from the drivers.
There was a problem hiding this comment.
thanks for this PR @LuukBlom!
I comfired it is working for zarr and tif, but not yet for nc files. See code snippet below for a test.
It would be good to also add some examples to the docs as there are some specificities that are good for users to know. Like nc only works with the "h5netcdf" engine, not with "netcdf4".
I also tested for vector files where I think another issue (#1542) related to MinIO specific (or different endpoint) seems to be the issue. I made a seperate issue for that one as I can imagine it requires a new PR.
import s3fs
import xarray as xr
from hydromt import DataCatalog
s3_path = "s3://hydromt-data/test/chirps-v2.0.1981.days_p05.nc"
## test with Hydromt DataCatalog
dc = DataCatalog().from_dict(
{
"test": {
"data_type": "RasterDataset",
"uri": s3_path,
"driver": {
"name": "raster_xarray",
"filesystem": {
"protocol": "s3",
"anon": False,
"profile": "hydromt-data",
},
# NOTE: requires h5netcdf engine
"options": {"engine": "h5netcdf"},
},
},
}
)
try:
ds = dc.get_rasterdataset("test")
print(ds)
except Exception as e:
print(e)
# > PermissionError: Forbidden
## test with s3fs and xarray (this works)
fs = s3fs.S3FileSystem(anon=False, profile="hydromt-data")
with fs.open(s3_path, mode="rb") as s3_file:
ds = xr.open_dataset(s3_file, engine="h5netcdf")
print(ds)…urces. TODO: add docs for how to setup aws profile (see hydromt_data_pipelines).
ClaireDons
left a comment
There was a problem hiding this comment.
Hey, I only managed to partly review it for now and mainly looked at the docs, which look mostly fine. Although in the future it might be nice to have a full example in the working with models section that gets data from an S3 bucket. @JoerivanEngelen will review the rest!
JoerivanEngelen
left a comment
There was a problem hiding this comment.
Looks good. I like the generalization of the readers.
I have some small comments though.
DirkEilander
left a comment
There was a problem hiding this comment.
I've tested that vector files work. Found one small potential issue in the docs.
Will try to also test nc later today.
NC and tif files also work in my tests. This PR can be merged I think. Thanks @LuukBlom |
This PR adds filesystem-aware reading across HydroMT readers and improves support for private cloud storage.
Note that the vast majority of lines changed are in
pixi.lock, added testing, refactoring the readers, and updating callers of the readers.Issue addressed
Fixes #1530
Fixes #1542
Explanation
Took Dirks changes as a starting point.
hydromt.data_catalog.drivers.xarray_options._read_xarraytests/data_catalog/test_private_bucket.pythat reads from the (now still) private hydromt-data minio bucketopen_zarrs&open_mfdatasettohydromt.readersEdit:
After deciding to add AbstractFilesystem to the args for all readers:
hydromt._fsio.pyfor opening handles, reading bytes, creating Zarr mappers, closing lazy resources, and normalizing permission errors.hydromt.readers.py: raster, vector, tabular, NetCDF, Zarr, and multi-file readers to accept fsspec filesystems.storage_optionsonce at the filesystem boundary.Most relevant files to review:
hydromt/_fsio.py: Shared fsspec utilities for remote readeinghydromt/readers.py: Reader implementations updated to consume filesystemstests/test_readers_filesystem.py: In-memory filesystem tests covering remote reads, lazy loadingtests/data_catalog/test_private_bucket.py: Private S3 integration and permission-handling tests