Skip to content

Seed catalogs for twelve South Asian languages - #1678

Merged
dqnykamp merged 5 commits into
Doenet:mainfrom
dqnykamp:seed-south-asian-catalogs
Aug 10, 2026
Merged

Seed catalogs for twelve South Asian languages#1678
dqnykamp merged 5 commits into
Doenet:mainfrom
dqnykamp:seed-south-asian-catalogs

Conversation

@dqnykamp

@dqnykamp dqnykamp commented Aug 10, 2026

Copy link
Copy Markdown
Member

Add message catalogs for Sanskrit, Maithili, Bhojpuri, Konkani, Dogri, Bodo, Manipuri, Santali, Kashmiri, Dhivehi, Tibetan and Dzongkha.

Each covers all four namespaces — the viewer chrome, the editor and language-server surfaces, the prose the core computes into a document, and the warnings and errors. documentLocale="sa" and <document lang="dz"> work with nothing configured, and all twelve reach <document lang>'s autocomplete.

These are unreviewed machine-generated seeds, and every file says so in its header. Nothing falls back silently: a key a translation is missing renders in English, which is what makes seeding safe.

The finding: a script says nothing about the fork

The batch adds five Devanagari catalogs that answer the agreement question three different ways. Sanskrit selects on $gender and $role; Konkani selects on both, with Marathi's three genders; Dogri selects on $gender alone; Maithili and Bhojpuri select on neither. Hindi, Marathi and Nepali were already three more answers in the same letters. styleDescriptions.test.ts pins all five side by side so the claim cannot quietly stop being true.

Sanskrit hands each clause position a different case — nominative standing alone, instrumental before «सीमया सह», locative before «पृष्ठभूमौ», nominative neuter agreeing with «पाठ्यम्». That is the shape de, ru, pl, cs, hr, sr and sl already have, three genders and all four positions; what is new is which case each position governs, and that sa writes the fork in all sixteen adjectives where kok writes it in nine. The three clause branches are one form each rather than three, because the noun each lands on is a word the catalog writes, so $gender is not consulted inside them. Sandhi is the affix rule at a word boundary, and the way out is the one Sanskrit itself supplies: the pada-pāṭha, unsandhied across every placeable and applied only where both words are the catalog's own.

Dogri is the one worth reading closely, because it is not a claim that Dogri does not inflect. Its masculine -आ adjectives take an oblique, and none of the three clause positions reaches it — the border and background are feminine and the text colour is a direct masculine. A $role branch would render what the $gender branch underneath already renders. Its header records which noun would have to change for the answer to change, and the test asserts काली रेखा against काला बिंदू so that a future noun entry in a masculine oblique position would fail.

Plural categories at both extremes

Santali writes a [two] branch in every counted message it has — 16 of them, more than any other catalog, ahead of Slovene's 14, Arabic's ten and Hebrew's eight. Intl.PluralRules("sat") reports one, two and other, because Santali marks a dual with -ᱠᱤᱱ where the plural takes -ᱠᱚ. Nine catalogs already resolve two; none writes it everywhere.

Tibetan and Dzongkha are the opposite extreme: ICU reports exactly one category for each, so no message in either can select on a count, and every counted message in both is written flat — the shape locales/ja, locales/th and the other one-category catalogs already use. The [0] branches that survive are matched by number rather than by category — Fluent resolves an explicit number before it consults the plural rules, so a wording for none is still reachable.

Tibetan case particles, and where the way out runs out

A Tibetan case particle is chosen by the final letter of the syllable before it: the agentive is གིས་, ཀྱིས་, གྱིས་ or ཡིས་ and the genitive གི་, ཀྱི་, གྱི་ or ཡི་. Beside a placeable there is no syllable to look at. So bo/content.ftl and dz/content.ftl — the two files that compose a phrase, and so the two that get to choose their particles — use only the invariant ones: དང་, ལ་ (Dzongkha ལུ་), ནང་. That is prefer the free allomorph over the bound one applied to a particle rather than to a prefix, and the third family to take that way out after the Bisayan catalogs and Kʼicheʼ. bo and dz get a row of their own in the affix table.

The other three files in each are where that way out runs out, and this is the first batch to show the boundary. A message that names a thing the document contains — summary-statistics-caption, the fifteen variant-* diagnostics — needs a genitive, and there is no invariant genitive to reach for. Both catalogs write the default shape, གི་ and གིས་, everywhere: right after most syllables and wrong after some. A single wrong form is findable where a scatter of them is not, so the shape is uniform on purpose and both headers say so.

Word order, and what decides the $part split

Three catalogs put their adjectives after the noun — Meitei and both Tibetan-script ones — and all three therefore reach [noun-tail] for the regular polygon, because a side count is a complement there. That is the Austronesian batch's lesson restated from a different family: what decides the split is the shape of the complement, not the side the adjectives sit on. The nine prenominal catalogs fold the count into the head and leave the tail empty.

The chemistry gap splits five ways

All twelve leave element-name and element-anion-name to English, and the README now carries a paragraph for the batch rather than only twelve more names in the roster list. Seven are the school-system case in six different school systems. Bodo and Sanskrit are the Samoan case: Bodo-medium schooling in Assam does run to the secondary grades and Sanskrit has words for the metals it knew, but neither has a settled list of all 118. Santali is the Ojibwe shape, schooling that stops below the table and no table behind it; Dhivehi is locales/to's, both halves at once. Tibetan alone is the Khmer case — the names exist and no single convention does — while dz, in the same script, is the plain school-system case: two catalogs in one script, two opposite reasons for the same gap.

Two limits recorded rather than worked around

locales/ks writes every adjective in the masculine citation form although Kashmiri agrees them for gender. noun-gender is filled in with the real genders anyway, so a speaker adding the fork has the table already and need only write the feminine forms. That is a deliberate gap, not a claim about the language.

locales/dv puts piecewise-condition-if on the wrong side of the mathematics: Dhivehi's conditional particle «ނަމަ» is clause-final and the renderer places that key before what it introduces. That is the locales/tpi shape — a distinction the composition messages do not expose — and splitting the key into a prefix and a suffix is a change to the worker that no existing catalog needs and this one would use.

One word cannot be three

mni would render “type”, “attribute” and “variant” all as মখল if each were translated by sense, and the editor puts two of them side by side in the same panel. মখল is kept for type; the other two take the words locales/bn and locales/as use in the same script — এট্রিবিউট for an attribute and রূপভেদ for a variant — and both headers say so. locales/sa had one label in the nominative plural where the same English word is singular elsewhere; both now read सूचना.

Scripts and naming

Three scripts are new: Ol Chiki (Santali), Thaana (Dhivehi) and Tibetan (bo and dz). Kashmiri is Perso-Arabic and Manipuri is Bengali, both already in the roster.

Six script asymmetries arrive with no new answer to any of them. sa is written in more scripts than anything else here, so sa-Gran, sa-Knda and a dozen more reach the Devanagari catalog; kok-Latn (Romi Konkani), doi-Arab, sat-Deva/sat-Beng/sat-Orya/sat-Latn and ks-Deva all reach theirs. mni is the one that will be argued with: the rule is that a catalog is written in whatever CLDR fills a bare tag in as — the rule that makes sr Cyrillic and az Latin — and mni maximizes to mni-Beng, so a reader arriving under mni-Mtei gets Bengali letters even though Meetei Mayek is what Manipur's schools now teach. Here more than anywhere the usual answer, a second catalog beside the first, is owed rather than hypothetical, and the header says so.

bo and dz are two directories rather than one with a script tag: the hr-against-sr case a fifth time, two standard languages and two vocabularies in one script. dz writes ཧོནམ where bo writes སྔོན་པོ, སྦོམ where it writes མཐུག་པོ, and ལུ་ where it writes ལ་.

dv is the batch's locale whose endonym comes back as its English name, so the roster reads "Divehi" once — locales/co's case with Klingon's twist, since Dhivehi has a well-known endonym, «ދިވެހި», that CLDR does not carry.

Negotiation

kok and doi are ISO 639-3 macrolanguages and join MACROLANGUAGE_MEMBERS for the same published reason qu, oj and bik did. The catalogs are Goan Konkani and Dogri proper; knn (Maharashtrian Konkani) and xnr (Kangri) reach them, and gom and dgo are the members ICU already folds, listed anyway so each group reads as a whole. That takes the map to eight keys, seven of them macrolanguages, which its docstring now says. Nothing else needs an alias: san, bod, dzo and div are canonicalized by Intl.getCanonicalLocales before negotiation.

Tests

  • packages/i18n/test/negotiate.test.ts: both macrolanguage folds across four members, the ISO 639-3 codes ICU handles unaided, nine script tags and four region tags — and the negative controls kfy, mag, hoc, njz and grt, neighbours of mai, bho, sat and brx that belong to no macrolanguage with a catalog and are held on English. Dropping knn and xnr from the map fails it.
  • packages/i18n/test/direction.test.ts: ks and dv added to the written-out right-to-left list, which the two tests hold from opposite sides.
  • packages/utils/test/styleDescriptions.test.ts: Sanskrit's case in each clause position, Konkani's three genders and oblique, the three post-nominal catalogs and their [noun-tail] polygons asserted whole, and Dogri's gender fork pinned against the $role branch it does not need. The whole-string form matters: an earlier toContain("5") on the polygon passed just as happily against a catalog with an empty [tail], and the mai/bho substring check could not have seen an oblique appear, a direct form being a prefix of one.

Docs

packages/i18n/README.md gains the twelve codes in its roster, bs and yi added to the $role list they were already missing from, a "The South Asian batch" section, a chemistry paragraph for the batch, a bo/dz row in the affix table with the particle case written up beneath it, kok and sa added to the list of catalogs that select on $role, and the right-to-left section taken from eight ships to ten with ks and dv in its table. Three older claims were corrected against the current catalogs while checking the new ones: yi is not the only right-to-left catalog that forks on $role (ur, ps and sd do too), sa's fork is the widest the mechanism allows rather than wider than any other catalog's, and the free-allomorph way out is now taken by three families rather than three catalogs.

Regenerated

supportedLocales.ts (160 locales) and doenet-schema.json, via codegen → build i18n → build:schema → build static-assets.

🤖 Generated with Claude Code

https://claude.ai/code/session_01BnWCnDrhzdXPBDvjeVuXVX

dqnykamp and others added 5 commits August 9, 2026 22:15
Sanskrit, Maithili, Bhojpuri, Konkani, Dogri, Bodo, Manipuri, Santali,
Kashmiri, Dhivehi, Tibetan and Dzongkha. All twelve are unreviewed
machine-generated seeds, and every file says so in its header.

Konkani and Dogri join MACROLANGUAGE_MEMBERS, so `knn` and `xnr` reach
their catalogs. Kashmiri and Dhivehi bring the roster's right-to-left
catalogs to ten and need nothing from direction.ts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BnWCnDrhzdXPBDvjeVuXVX
… them

Review pass over the twelve seeds. The catalogs were structurally sound —
lint passes, no key is absent from `en`, no placeable or DoenetML identifier
was dropped or translated — so nearly all of this is prose that asserted more
than the files did.

Claims corrected:

- Santali was not the first non-Semitic `two`: nine catalogs already resolve
  and write it, six of them non-Semitic. What is true is that `sat` writes it
  in every counted message it has, which is now what the header and README
  say. Its two all-identical selects are written flat, as `he` does.
- `bo` and `dz` claim to use only invariant case particles. That holds in
  `content.ftl` and cannot hold in the other three, which need a genitive
  after a placeable in 45 messages. The limit is now recorded, and `bo`'s
  three stray allomorphs are normalized to the same default the rest write.
- `mni` claimed twice to be the only post-nominal catalog in the batch, with
  `bo`'s header naming it as one of three.
- Five headers claimed Indian digit grouping that CLDR does not give those
  locales; `sa`, `kok`, `brx` and `dz` really do get it.
- `sa` claimed the widest `$role` fork and a number inflection it never
  writes; `mai` and `bho` listed their invariant adjectives and left out the
  two that would test the claim; `ks/diagnostics.ftl` said its branches read
  alike where they differ; `brx` miscounted its own colour prefix.
- The roster list is merged alphabetically, as every earlier batch did, and
  the right-to-left section's stale "the eight" follows "Ten ship".

Content fixed: `kok` gave आयत a gender Marathi does not and used «भितर»
where the sense is partitive; `brx` rendered Part and Section, and Problem
and complexity, with one word each; `mni` collided Section with Part; `dz`
wrote a Tibetan dative beside its own; `bo` and `dz` had four degenerate
single-variant selects; `bho` broke its own imperative rule in one label.

Konkani is added to `styleDescriptions.test.ts`, which the README said
pinned all five Devanagari catalogs and pinned four.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BnWCnDrhzdXPBDvjeVuXVX
…aded word

`bo` and `dz` have a single plural category, and both headers said every
counted message was written flat — but twelve messages kept a `{ $count ->
*[other] … }` wrapper with nothing to choose between. Every other
one-category catalog (`ja`, `th`, `km`) writes those flat, so these do now
and the headers describe what the files contain.

`locales/mni` rendered "type", "attribute" and "variant" all as মখল, two of
them in the same editor panel. মখল now means type and nothing else; an
attribute is এট্রিবিউট and a variant রূপভেদ, the words `locales/bn` and
`locales/as` use in the same script.

Also: `locales/sa`'s Info tab label was a nominative plural where the same
English word is singular in the heading beside it; the README's `[two]`
comparison skipped Arabic's ten branches; and its `$role` list had never
included `bs` or `yi`, both of which fork on it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BnWCnDrhzdXPBDvjeVuXVX
…o fail

The `[noun-tail]` assertion passed against a catalog with an empty `[tail]`:
emptying `dz`'s tail and folding the count into `[head]` left `toContain("5")`
true and `endsWith("5")` false either way, so it pinned nothing. All three
post-nominal outputs are now asserted whole, and so are `mai`'s and `bho`'s,
where a substring check could not have seen an oblique appear — a direct form
is a prefix of one.

`MACROLANGUAGE_MEMBERS`'s docstring still said "five of the six keys" after
`kok` and `doi` made it eight, and did not list `gom` and `dgo` among the
members CLDR folds on its own.

The README's new section pointed at a chemistry paragraph for the batch that
was never written; the twelve had only been added to the roster list. Written
now, and it splits five ways rather than the three the changeset claimed —
Bodo and Sanskrit are the Samoan case, Santali the Ojibwe one and Dhivehi
`locales/to`'s, none of which is the school-system case the other seven are.

Three claims did not survive checking: `locales/sa` does not select on "more
of each than any of them" (`de`, `ru`, `pl`, `cs`, `hr`, `sr` and `sl` all have
three genders and the same four positions, and `kok` has them too), the batch
enumerates six script asymmetries rather than five, and `yi` is not the only
right-to-left catalog that forks on `$role` — `ur`, `ps` and `sd` all do, as
the same file's own list says. `bo` and `dz` also now have a row in the affix
table the paragraph beneath calls their entry in.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BnWCnDrhzdXPBDvjeVuXVX
Cosmetic only: an earlier edit left a mid-paragraph short line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BnWCnDrhzdXPBDvjeVuXVX
@dqnykamp
dqnykamp merged commit 9f2a2e1 into Doenet:main Aug 10, 2026
24 checks passed
@dqnykamp
dqnykamp deleted the seed-south-asian-catalogs branch August 10, 2026 07:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant