Seed catalogs for twelve South Asian languages - #1678
Merged
Conversation
Sanskrit, Maithili, Bhojpuri, Konkani, Dogri, Bodo, Manipuri, Santali, Kashmiri, Dhivehi, Tibetan and Dzongkha. All twelve are unreviewed machine-generated seeds, and every file says so in its header. Konkani and Dogri join MACROLANGUAGE_MEMBERS, so `knn` and `xnr` reach their catalogs. Kashmiri and Dhivehi bring the roster's right-to-left catalogs to ten and need nothing from direction.ts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BnWCnDrhzdXPBDvjeVuXVX
… them Review pass over the twelve seeds. The catalogs were structurally sound — lint passes, no key is absent from `en`, no placeable or DoenetML identifier was dropped or translated — so nearly all of this is prose that asserted more than the files did. Claims corrected: - Santali was not the first non-Semitic `two`: nine catalogs already resolve and write it, six of them non-Semitic. What is true is that `sat` writes it in every counted message it has, which is now what the header and README say. Its two all-identical selects are written flat, as `he` does. - `bo` and `dz` claim to use only invariant case particles. That holds in `content.ftl` and cannot hold in the other three, which need a genitive after a placeable in 45 messages. The limit is now recorded, and `bo`'s three stray allomorphs are normalized to the same default the rest write. - `mni` claimed twice to be the only post-nominal catalog in the batch, with `bo`'s header naming it as one of three. - Five headers claimed Indian digit grouping that CLDR does not give those locales; `sa`, `kok`, `brx` and `dz` really do get it. - `sa` claimed the widest `$role` fork and a number inflection it never writes; `mai` and `bho` listed their invariant adjectives and left out the two that would test the claim; `ks/diagnostics.ftl` said its branches read alike where they differ; `brx` miscounted its own colour prefix. - The roster list is merged alphabetically, as every earlier batch did, and the right-to-left section's stale "the eight" follows "Ten ship". Content fixed: `kok` gave आयत a gender Marathi does not and used «भितर» where the sense is partitive; `brx` rendered Part and Section, and Problem and complexity, with one word each; `mni` collided Section with Part; `dz` wrote a Tibetan dative beside its own; `bo` and `dz` had four degenerate single-variant selects; `bho` broke its own imperative rule in one label. Konkani is added to `styleDescriptions.test.ts`, which the README said pinned all five Devanagari catalogs and pinned four. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BnWCnDrhzdXPBDvjeVuXVX
…aded word
`bo` and `dz` have a single plural category, and both headers said every
counted message was written flat — but twelve messages kept a `{ $count ->
*[other] … }` wrapper with nothing to choose between. Every other
one-category catalog (`ja`, `th`, `km`) writes those flat, so these do now
and the headers describe what the files contain.
`locales/mni` rendered "type", "attribute" and "variant" all as মখল, two of
them in the same editor panel. মখল now means type and nothing else; an
attribute is এট্রিবিউট and a variant রূপভেদ, the words `locales/bn` and
`locales/as` use in the same script.
Also: `locales/sa`'s Info tab label was a nominative plural where the same
English word is singular in the heading beside it; the README's `[two]`
comparison skipped Arabic's ten branches; and its `$role` list had never
included `bs` or `yi`, both of which fork on it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BnWCnDrhzdXPBDvjeVuXVX
…o fail
The `[noun-tail]` assertion passed against a catalog with an empty `[tail]`:
emptying `dz`'s tail and folding the count into `[head]` left `toContain("5")`
true and `endsWith("5")` false either way, so it pinned nothing. All three
post-nominal outputs are now asserted whole, and so are `mai`'s and `bho`'s,
where a substring check could not have seen an oblique appear — a direct form
is a prefix of one.
`MACROLANGUAGE_MEMBERS`'s docstring still said "five of the six keys" after
`kok` and `doi` made it eight, and did not list `gom` and `dgo` among the
members CLDR folds on its own.
The README's new section pointed at a chemistry paragraph for the batch that
was never written; the twelve had only been added to the roster list. Written
now, and it splits five ways rather than the three the changeset claimed —
Bodo and Sanskrit are the Samoan case, Santali the Ojibwe one and Dhivehi
`locales/to`'s, none of which is the school-system case the other seven are.
Three claims did not survive checking: `locales/sa` does not select on "more
of each than any of them" (`de`, `ru`, `pl`, `cs`, `hr`, `sr` and `sl` all have
three genders and the same four positions, and `kok` has them too), the batch
enumerates six script asymmetries rather than five, and `yi` is not the only
right-to-left catalog that forks on `$role` — `ur`, `ps` and `sd` all do, as
the same file's own list says. `bo` and `dz` also now have a row in the affix
table the paragraph beneath calls their entry in.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BnWCnDrhzdXPBDvjeVuXVX
Cosmetic only: an earlier edit left a mid-paragraph short line. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BnWCnDrhzdXPBDvjeVuXVX
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Add message catalogs for Sanskrit, Maithili, Bhojpuri, Konkani, Dogri, Bodo, Manipuri, Santali, Kashmiri, Dhivehi, Tibetan and Dzongkha.
Each covers all four namespaces — the viewer chrome, the editor and language-server surfaces, the prose the core computes into a document, and the warnings and errors.
documentLocale="sa"and<document lang="dz">work with nothing configured, and all twelve reach<document lang>'s autocomplete.These are unreviewed machine-generated seeds, and every file says so in its header. Nothing falls back silently: a key a translation is missing renders in English, which is what makes seeding safe.
The finding: a script says nothing about the fork
The batch adds five Devanagari catalogs that answer the agreement question three different ways. Sanskrit selects on
$genderand$role; Konkani selects on both, with Marathi's three genders; Dogri selects on$genderalone; Maithili and Bhojpuri select on neither. Hindi, Marathi and Nepali were already three more answers in the same letters.styleDescriptions.test.tspins all five side by side so the claim cannot quietly stop being true.Sanskrit hands each clause position a different case — nominative standing alone, instrumental before «सीमया सह», locative before «पृष्ठभूमौ», nominative neuter agreeing with «पाठ्यम्». That is the shape
de,ru,pl,cs,hr,srandslalready have, three genders and all four positions; what is new is which case each position governs, and thatsawrites the fork in all sixteen adjectives wherekokwrites it in nine. The three clause branches are one form each rather than three, because the noun each lands on is a word the catalog writes, so$genderis not consulted inside them. Sandhi is the affix rule at a word boundary, and the way out is the one Sanskrit itself supplies: the pada-pāṭha, unsandhied across every placeable and applied only where both words are the catalog's own.Dogri is the one worth reading closely, because it is not a claim that Dogri does not inflect. Its masculine -आ adjectives take an oblique, and none of the three clause positions reaches it — the border and background are feminine and the text colour is a direct masculine. A
$rolebranch would render what the$genderbranch underneath already renders. Its header records which noun would have to change for the answer to change, and the test asserts काली रेखा against काला बिंदू so that a futurenounentry in a masculine oblique position would fail.Plural categories at both extremes
Santali writes a
[two]branch in every counted message it has — 16 of them, more than any other catalog, ahead of Slovene's 14, Arabic's ten and Hebrew's eight.Intl.PluralRules("sat")reportsone,twoandother, because Santali marks a dual with -ᱠᱤᱱ where the plural takes -ᱠᱚ. Nine catalogs already resolvetwo; none writes it everywhere.Tibetan and Dzongkha are the opposite extreme: ICU reports exactly one category for each, so no message in either can select on a count, and every counted message in both is written flat — the shape
locales/ja,locales/thand the other one-category catalogs already use. The[0]branches that survive are matched by number rather than by category — Fluent resolves an explicit number before it consults the plural rules, so a wording for none is still reachable.Tibetan case particles, and where the way out runs out
A Tibetan case particle is chosen by the final letter of the syllable before it: the agentive is གིས་, ཀྱིས་, གྱིས་ or ཡིས་ and the genitive གི་, ཀྱི་, གྱི་ or ཡི་. Beside a placeable there is no syllable to look at. So
bo/content.ftlanddz/content.ftl— the two files that compose a phrase, and so the two that get to choose their particles — use only the invariant ones: དང་, ལ་ (Dzongkha ལུ་), ནང་. That is prefer the free allomorph over the bound one applied to a particle rather than to a prefix, and the third family to take that way out after the Bisayan catalogs and Kʼicheʼ.boanddzget a row of their own in the affix table.The other three files in each are where that way out runs out, and this is the first batch to show the boundary. A message that names a thing the document contains —
summary-statistics-caption, the fifteenvariant-*diagnostics — needs a genitive, and there is no invariant genitive to reach for. Both catalogs write the default shape, གི་ and གིས་, everywhere: right after most syllables and wrong after some. A single wrong form is findable where a scatter of them is not, so the shape is uniform on purpose and both headers say so.Word order, and what decides the
$partsplitThree catalogs put their adjectives after the noun — Meitei and both Tibetan-script ones — and all three therefore reach
[noun-tail]for the regular polygon, because a side count is a complement there. That is the Austronesian batch's lesson restated from a different family: what decides the split is the shape of the complement, not the side the adjectives sit on. The nine prenominal catalogs fold the count into the head and leave the tail empty.The chemistry gap splits five ways
All twelve leave
element-nameandelement-anion-nameto English, and the README now carries a paragraph for the batch rather than only twelve more names in the roster list. Seven are the school-system case in six different school systems. Bodo and Sanskrit are the Samoan case: Bodo-medium schooling in Assam does run to the secondary grades and Sanskrit has words for the metals it knew, but neither has a settled list of all 118. Santali is the Ojibwe shape, schooling that stops below the table and no table behind it; Dhivehi islocales/to's, both halves at once. Tibetan alone is the Khmer case — the names exist and no single convention does — whiledz, in the same script, is the plain school-system case: two catalogs in one script, two opposite reasons for the same gap.Two limits recorded rather than worked around
locales/kswrites every adjective in the masculine citation form although Kashmiri agrees them for gender.noun-genderis filled in with the real genders anyway, so a speaker adding the fork has the table already and need only write the feminine forms. That is a deliberate gap, not a claim about the language.locales/dvputspiecewise-condition-ifon the wrong side of the mathematics: Dhivehi's conditional particle «ނަމަ» is clause-final and the renderer places that key before what it introduces. That is thelocales/tpishape — a distinction the composition messages do not expose — and splitting the key into a prefix and a suffix is a change to the worker that no existing catalog needs and this one would use.One word cannot be three
mniwould render “type”, “attribute” and “variant” all as মখল if each were translated by sense, and the editor puts two of them side by side in the same panel. মখল is kept for type; the other two take the wordslocales/bnandlocales/asuse in the same script — এট্রিবিউট for an attribute and রূপভেদ for a variant — and both headers say so.locales/sahad one label in the nominative plural where the same English word is singular elsewhere; both now read सूचना.Scripts and naming
Three scripts are new: Ol Chiki (Santali), Thaana (Dhivehi) and Tibetan (
boanddz). Kashmiri is Perso-Arabic and Manipuri is Bengali, both already in the roster.Six script asymmetries arrive with no new answer to any of them.
sais written in more scripts than anything else here, sosa-Gran,sa-Kndaand a dozen more reach the Devanagari catalog;kok-Latn(Romi Konkani),doi-Arab,sat-Deva/sat-Beng/sat-Orya/sat-Latnandks-Devaall reach theirs.mniis the one that will be argued with: the rule is that a catalog is written in whatever CLDR fills a bare tag in as — the rule that makessrCyrillic andazLatin — andmnimaximizes tomni-Beng, so a reader arriving undermni-Mteigets Bengali letters even though Meetei Mayek is what Manipur's schools now teach. Here more than anywhere the usual answer, a second catalog beside the first, is owed rather than hypothetical, and the header says so.boanddzare two directories rather than one with a script tag: thehr-against-srcase a fifth time, two standard languages and two vocabularies in one script.dzwrites ཧོནམ wherebowrites སྔོན་པོ, སྦོམ where it writes མཐུག་པོ, and ལུ་ where it writes ལ་.dvis the batch's locale whose endonym comes back as its English name, so the roster reads "Divehi" once —locales/co's case with Klingon's twist, since Dhivehi has a well-known endonym, «ދިވެހި», that CLDR does not carry.Negotiation
kokanddoiare ISO 639-3 macrolanguages and joinMACROLANGUAGE_MEMBERSfor the same published reasonqu,ojandbikdid. The catalogs are Goan Konkani and Dogri proper;knn(Maharashtrian Konkani) andxnr(Kangri) reach them, andgomanddgoare the members ICU already folds, listed anyway so each group reads as a whole. That takes the map to eight keys, seven of them macrolanguages, which its docstring now says. Nothing else needs an alias:san,bod,dzoanddivare canonicalized byIntl.getCanonicalLocalesbefore negotiation.Tests
packages/i18n/test/negotiate.test.ts: both macrolanguage folds across four members, the ISO 639-3 codes ICU handles unaided, nine script tags and four region tags — and the negative controlskfy,mag,hoc,njzandgrt, neighbours ofmai,bho,satandbrxthat belong to no macrolanguage with a catalog and are held on English. Droppingknnandxnrfrom the map fails it.packages/i18n/test/direction.test.ts:ksanddvadded to the written-out right-to-left list, which the two tests hold from opposite sides.packages/utils/test/styleDescriptions.test.ts: Sanskrit's case in each clause position, Konkani's three genders and oblique, the three post-nominal catalogs and their[noun-tail]polygons asserted whole, and Dogri's gender fork pinned against the$rolebranch it does not need. The whole-string form matters: an earliertoContain("5")on the polygon passed just as happily against a catalog with an empty[tail], and themai/bhosubstring check could not have seen an oblique appear, a direct form being a prefix of one.Docs
packages/i18n/README.mdgains the twelve codes in its roster,bsandyiadded to the$rolelist they were already missing from, a "The South Asian batch" section, a chemistry paragraph for the batch, abo/dzrow in the affix table with the particle case written up beneath it,kokandsaadded to the list of catalogs that select on$role, and the right-to-left section taken from eight ships to ten withksanddvin its table. Three older claims were corrected against the current catalogs while checking the new ones:yiis not the only right-to-left catalog that forks on$role(ur,psandsddo too),sa's fork is the widest the mechanism allows rather than wider than any other catalog's, and the free-allomorph way out is now taken by three families rather than three catalogs.Regenerated
supportedLocales.ts(160 locales) anddoenet-schema.json, via codegen → build i18n → build:schema → build static-assets.🤖 Generated with Claude Code
https://claude.ai/code/session_01BnWCnDrhzdXPBDvjeVuXVX