Skip to content

Latest commit

 

History

History
83 lines (66 loc) · 3.03 KB

File metadata and controls

83 lines (66 loc) · 3.03 KB

Serving data from S3

The site never talks to a backend. Two things touch the bucket:

  1. The build (dataherb catalog build in CI) reads metadata and status files and lists prefixes. It uses boto3 with the normal AWS credential chain (in GitHub Actions: OIDC via AWS_ROLE_ARN, see the deploy workflows).
  2. The reader's browser downloads data files (previews, explorer) and latest.json status files (live status) over HTTPS.

Suggested layout

s3://my-company-datalake/
  datasets/
    orders-daily/
      dataherb.yml
      orders.parquet
    regions/
      dataherb.yml
      regions.csv
  _dataherb/status/
    orders-export/
      latest.json
      runs/20261007T020012Z-scheduled__2026-10-07.json

Build permissions

Read-only is enough:

{
  "Effect": "Allow",
  "Action": ["s3:GetObject", "s3:ListBucket"],
  "Resource": ["arn:aws:s3:::my-company-datalake", "arn:aws:s3:::my-company-datalake/*"]
}

Jobs that write status need s3:PutObject (and s3:GetObject to carry last_success forward) on _dataherb/status/*.

Browser access: pick one

Option How Good for
CloudFront in front of the bucket Origin Access Control; restrict the distribution with a WAF IP set (office/VPN ranges) or CloudFront signed cookies from your SSO. Set public_base_url to the distribution URL. Most companies. Fast, cacheable, no public bucket.
Bucket policy limited to the corporate network aws:SourceIp (VPN egress) or aws:SourceVpce (VPC endpoint) condition on s3:GetObject. Simple internal setups.
Presigned URLs presign: true on the store. The build signs every data URL. Small catalogs without a CDN. URLs expire (max 7 days); rebuild well within presign_expiry_hours. Live status cannot be presigned beyond that either.
Public bucket Public-read on datasets/ only. Genuinely public open data.

CORS

The explorer fetches files with JavaScript, so the bucket (or CloudFront response headers policy) must allow your site's origin and range requests:

[
  {
    "AllowedOrigins": ["https://data-catalog.example.com"],
    "AllowedMethods": ["GET", "HEAD"],
    "AllowedHeaders": ["Range", "If-None-Match", "Cache-Control"],
    "ExposeHeaders": ["Content-Length", "Content-Range", "Content-Type", "ETag", "Accept-Ranges"],
    "MaxAgeSeconds": 3600
  }
]
aws s3api put-bucket-cors --bucket my-company-datalake --cors-configuration file://cors.json

Range and the exposed Content-Range let DuckDB read only the parts of Parquet files a query needs. Without them it falls back to downloading whole files (you will see "fall back to full HTTP read" in the console).

Hosting the site itself on S3

Set the repository variables DEPLOY_TARGET=s3, SITE_BUCKET, AWS_ROLE_ARN, AWS_REGION and optionally CLOUDFRONT_DISTRIBUTION_ID; .github/workflows/deploy-s3.yml then publishes dist/ on every build. The site is plain files with hash-based routes, so any static host works (S3 website, CloudFront, GitHub Pages, Netlify, nginx).