Repository navigation
Add the dataset-* repos; catalog entries as Markdown - #1
Merged
Merged
Conversation
Generated with `dataherb catalog add --match dataset` for the dataset-* repos in the DataHerb org, then named and described by hand. amazon-categories and economic-trading-indicators are private and need the CATALOG_GITHUB_TOKEN secret to resolve at build time. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ggg4AutCcRdK8c57NMxSqp
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ggg4AutCcRdK8c57NMxSqp
Every catalog/*.yml becomes catalog/*.md: the same fields in front matter, plus a body rendered on the dataset page above the dataset's own documentation. The new dataset-* entries get bodies with sources, file lists and notes. The site's Markdown renderer now handles numbered lists, tables, block quotes, rules, images and bare URLs. dataherb is pinned to the dataherb-python#27 branch, which reads .md entries, until that PR is merged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ggg4AutCcRdK8c57NMxSqp
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011e9fA4TL7vcAAy9nC68yCH
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Requested by L · project thread
Before: the catalog listed 3 of the 7
dataset-*repos in the DataHerb org, and entries were plain YAML with no place for free text.After: all 7 repos are listed. Every entry is now
catalog/<id>.md, with the fields in YAML front matter and a Markdown body that the dataset page shows above the dataset's own documentation.New entries:
airpassengerentso-e-power-consumptionamazon-categorieseconomic-trading-indicatorsThese were already present and were skipped: covid-19-csse, data-science-jobs and samsung-device-updates.
How:
dataherb catalog add. Names, descriptions, tags, column descriptions and Markdown bodies (sources, file lists, notes) were written by hand..ymlentry was wrapped in front matter, so its fields and comments are unchanged.site/assets/lib/util.jshas a fuller Markdown renderer, still escaped and dependency-free. It adds numbered lists, tables, block quotes, rules, images and bare URLs, andstyle.csshas matching styles.dataherb catalog add.Dependency: DataHerb/dataherb-python#27 (reads
.mdentries) is merged, andpyproject.toml/uv.locknow pin dataherb back tomaster(d1d46b0).The two private repos resolve only when the
CATALOG_GITHUB_TOKENsecret holds a token that can read them. Without it, the build shows them with a "metadata unavailable" error. Browser previews of private raw files won't load either way, because raw.githubusercontent.com needs auth.🤖 Generated with Claude Code
https://claude.ai/code/session_01Ggg4AutCcRdK8c57NMxSqp