Skip to content

Add catalog and job status commands for DataHerb Explorer - #26

Merged
emptymalei merged 2 commits into
masterfrom
catalog-v2
Oct 7, 2026
Merged

emptymalei merged 2 commits into
masterfrom
catalog-v2

Conversation

@emptymalei

@emptymalei emptymalei commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Requested by L · project thread

Before: dataherb could create, validate, upload and serve git-hosted datasets, but there was no way to build the new static catalog site (DataHerb Explorer) or to monitor the jobs that refresh datasets.

After: the same CLI builds and checks the explorer catalog and writes and checks job status files:

Command
dataherb catalog build [-c dataherb.config.yml] [-o dist] [--site site] [--strict] Build the static explorer from the config, catalog entries and job status.
dataherb catalog validate Validate the config and catalog entries against JSON Schemas.
dataherb catalog lint [--min-score N] Metadata quality score (0 to 100) per dataset.
dataherb catalog serve [dist] Serve a built catalog locally.
dataherb status emit --target s3://... --job-id X --status running/success/failed ... Write a dataherb.status/v1 status file.
dataherb status check [--fail-on failing,stuck,stale] Print job health; exit 1 if any job is unhealthy.

dataherb create now drafts a v2 dataherb.json from the data files in a folder (columns, types, row counts), asks for the source (git, s3, http, local), owner, update frequency and status job, and has --no-input for scripts. dataherb validate adds a schema check and prints the metadata quality score. v1 dataherb.json and the legacy .dataherb/metadata.yml still parse.

How: a new dataherb.catalog package (config, stores, metadata resolution, status spec, lint, schemas) with click commands in dataherb/cmd/catalog.py. The site template itself lives in DataHerb/dataherb-explorer; catalog build uses site/ next to the config. New deps: PyYAML, jsonschema; extras s3 (boto3) and infer (duckdb). Also fixes dataherb upload using the cwd captured at import time, and replaces distutils.copy_tree so the CLI imports on Python 3.12. Tests in tests/catalog/; tutorials added to the docs.

🤖 Generated with Claude Code

https://claude.ai/code/session_016Zz2NMq7okdD9v71q6vtDq


Generated by Claude Code


Generated by Claude Code

Adds `dataherb catalog build|validate|lint|serve` and
`dataherb status emit|check`, backed by a new dataherb.catalog package
(config, stores for git/s3/http/local, metadata resolution, JSON Schemas,
the dataherb.status/v1 job status spec, metadata quality lint).

`dataherb create` now scaffolds a v2 dataherb.json from the data files
(with --no-input for scripts), and `dataherb validate` adds a schema check
and a metadata quality score. `dataherb upload` uses the current folder
instead of the import-time cwd.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016Zz2NMq7okdD9v71q6vtDq
@emptymalei emptymalei self-assigned this Oct 7, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016Zz2NMq7okdD9v71q6vtDq
emptymalei pushed a commit to DataHerb/dataherb-explorer that referenced this pull request Oct 7, 2026
The builder, validator, linter and status emitter moved into the existing
dataherb CLI (DataHerb/dataherb-python#26) as `dataherb catalog ...` and
`dataherb status ...`. This repo now keeps only the site, the config, the
catalog entries and the workflows. requirements.txt pins the dataherb
branch until that PR is released.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016Zz2NMq7okdD9v71q6vtDq
@emptymalei
emptymalei merged commit ac7f613 into master Oct 7, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants