Skip to content

feat: [50 MRG] Dataset index: public learner corpora with licenses (Closes #20) - #117

Open
laurentketterle-hub wants to merge 1 commit into
mergeos-bounties:masterfrom
laurentketterle-hub:feat/datasets-index
Open

feat: [50 MRG] Dataset index: public learner corpora with licenses (Closes #20)#117
laurentketterle-hub wants to merge 1 commit into
mergeos-bounties:masterfrom
laurentketterle-hub:feat/datasets-index

Conversation

@laurentketterle-hub

Copy link
Copy Markdown

Closes #20

Adds docs/DATASETS.md with 12 public ESL/learner corpora indexed by language, size, description, license, and link.

Content

  • 12 datasets covering EN, JA, KO, ZH, ES, DE, IT, CS, VI, and multi-language
  • Each entry includes: language codes, size, description, license, and official link
  • Usage notes section covering licensing, attribution, and calibration recommendations

Verification

cat docs/DATASETS.md | grep '^|' | wc -l  # 14 table rows (header + separator + 12 datasets)

@laurentketterle-hub

Copy link
Copy Markdown
Author

/attempt 20

@laurentketterle-hub

Copy link
Copy Markdown
Author

Friendly ping 👋 — this PR has been open for 3 days. Is there anything needed for review? Thanks!

@laurentketterle-hub

Copy link
Copy Markdown
Author

👋 Friendly nudge — this PR has been open for a while. Anything blocking merge, or happy to adjust if review feedback is needed? Thanks!

@laurentketterle-hub

Copy link
Copy Markdown
Author

/claim

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[50 MRG] Dataset index: public learner corpora with licenses

1 participant