The site never talks to a backend. Two things touch the bucket:
- The build (
dataherb catalog buildin CI) reads metadata and status files and lists prefixes. It uses boto3 with the normal AWS credential chain (in GitHub Actions: OIDC viaAWS_ROLE_ARN, see the deploy workflows). - The reader's browser downloads data files (previews, explorer) and
latest.jsonstatus files (live status) over HTTPS.
s3://my-company-datalake/
datasets/
orders-daily/
dataherb.yml
orders.parquet
regions/
dataherb.yml
regions.csv
_dataherb/status/
orders-export/
latest.json
runs/20261007T020012Z-scheduled__2026-10-07.json
Read-only is enough:
{
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:ListBucket"],
"Resource": ["arn:aws:s3:::my-company-datalake", "arn:aws:s3:::my-company-datalake/*"]
}Jobs that write status need s3:PutObject (and s3:GetObject to carry
last_success forward) on _dataherb/status/*.
| Option | How | Good for |
|---|---|---|
| CloudFront in front of the bucket | Origin Access Control; restrict the distribution with a WAF IP set (office/VPN ranges) or CloudFront signed cookies from your SSO. Set public_base_url to the distribution URL. |
Most companies. Fast, cacheable, no public bucket. |
| Bucket policy limited to the corporate network | aws:SourceIp (VPN egress) or aws:SourceVpce (VPC endpoint) condition on s3:GetObject. |
Simple internal setups. |
| Presigned URLs | presign: true on the store. The build signs every data URL. |
Small catalogs without a CDN. URLs expire (max 7 days); rebuild well within presign_expiry_hours. Live status cannot be presigned beyond that either. |
| Public bucket | Public-read on datasets/ only. |
Genuinely public open data. |
The explorer fetches files with JavaScript, so the bucket (or CloudFront response headers policy) must allow your site's origin and range requests:
[
{
"AllowedOrigins": ["https://data-catalog.example.com"],
"AllowedMethods": ["GET", "HEAD"],
"AllowedHeaders": ["Range", "If-None-Match", "Cache-Control"],
"ExposeHeaders": ["Content-Length", "Content-Range", "Content-Type", "ETag", "Accept-Ranges"],
"MaxAgeSeconds": 3600
}
]aws s3api put-bucket-cors --bucket my-company-datalake --cors-configuration file://cors.jsonRange and the exposed Content-Range let DuckDB read only the parts of
Parquet files a query needs. Without them it falls back to downloading whole
files (you will see "fall back to full HTTP read" in the console).
Set the repository variables DEPLOY_TARGET=s3, SITE_BUCKET,
AWS_ROLE_ARN, AWS_REGION and optionally CLOUDFRONT_DISTRIBUTION_ID;
.github/workflows/deploy-s3.yml then publishes dist/ on every build.
The site is plain files with hash-based routes, so any static host works
(S3 website, CloudFront, GitHub Pages, Netlify, nginx).