Developer scripts for maintaining Vernacula localisation, help files, and ONNX model exports.
The locale and help scripts require the anthropic Python package and an Anthropic API key:
pip install anthropic
export ANTHROPIC_API_KEY=sk-ant-... # or pass --api-key on each commandTo prevent format drift across the per-export directories, two scripts at
the scripts/ root are the canonical entry points after a successful
export run:
make_manifest.py— builds themanifest.jsonshape that Vernacula.Avalonia's download verifier reads ({"files": {<name>: {"md5": <hex>}}}). Backed by_export_utils/manifest.py.upload_to_hf.py— generic HuggingFace uploader, defaults to syncing the model card fromhf_readmes/<repo-basename>/README.md.
Typical post-export workflow:
python scripts/make_manifest.py --model-dir ~/models/<bundle> --all
python scripts/upload_to_hf.py \
--model-dir ~/models/<bundle> \
--repo-id christopherthompson81/<repo-name> \
--sync-readme --create-repoBoth take --exclude GLOB... so a bundle directory that has been run
from (ONNX Runtime writes *.ort / *.use-ort optimisation caches beside
the graphs) can still be published with --all:
python scripts/make_manifest.py --model-dir ~/models/kokoro --all --exclude '*.ort' '*.use-ort'
python scripts/upload_to_hf.py --model-dir ~/models/kokoro \
--repo-id christopherthompson81/kokoro-82m-onnx \
--exclude '*.ort' '*.use-ort' --sync-readme --create-repoThe C# contract is locked at
Vernacula.Avalonia/Services/ModelManagerService.cs::ParseManifestHashes
— don't fork the manifest shape per-export. New shared utilities go in
_export_utils/.
Translates new or changed UI string keys from en.json into all 24 supported language JSON files.
Before running, edit the two constants near the top of the file:
NEW_KEYS— the keys to add, in insertion order, with their English valuesANCHOR_KEY— the existing key after which the new keys will be inserted
The script only retranslates keys whose English source has changed since the last run (SQLite cache at src/Vernacula.Avalonia/Locales/locale_keys.db). Interrupted runs resume automatically.
# Translate any new/changed keys and write them into the locale JSON files
python scripts/update_locale_keys.py
# Force retranslation of all keys in NEW_KEYS, ignoring the cache
python scripts/update_locale_keys.py --force
# Pass the API key directly instead of via environment variable
python scripts/update_locale_keys.py --api-key sk-ant-...Translates new or changed English help Markdown files (Help/en/**/*.md) into all 24 supported language directories (Help/{lang}/).
Change detection is based on SHA-256 of the English file contents (cache at src/Vernacula.Avalonia/Help/help_files.db). The prompt instructs the model to preserve YAML frontmatter structure, Markdown syntax, relative link hrefs, backtick UI labels, and technical terms — only human-readable prose is translated.
# Translate all changed files for all languages
python scripts/update_help_files.py
# Force retranslation of everything, ignoring the cache
python scripts/update_help_files.py --force
# Retranslate one specific file for all languages
python scripts/update_help_files.py --files operations/editing_transcripts.md
# Process multiple specific files (comma-separated, relative to Help/en/)
python scripts/update_help_files.py --files index.md,operations/editing_transcripts.md
# Target only specific languages
python scripts/update_help_files.py --lang de,fr
# Combine filters — one file, two languages
python scripts/update_help_files.py --files index.md --lang bg,csScripts for exporting NeMo model checkpoints to ONNX. See nemo_export/README.md for full documentation.
export_parakeet_nemo_to_onnx.py— exports a Parakeet RNNT/TDT.nemocheckpoint to the ONNX package used by Vernaculaexport_sortformer_nemo_to_onnx.py— exports a streaming Sortformer diarization.nemocheckpoint to ONNXexport_silero_vad_to_onnx.py— exports Silero VAD to ONNXsetup_nemo_export_env.py— creates the Python export venvrequirements.txt— export dependencies
Exports DeepFilterNet3 as three streaming ONNX models with explicit GRU hidden-state I/O for chunk-by-chunk C# inference. See deepfilternet3_export/README.md for full documentation.
Exports the DiariZen diarization pipeline: segmentation model, WeSpeaker embedding model, and LDA/PLDA transform parameters. See diarizen_export/README.md for full documentation.
Exports CohereLabs/cohere-transcribe-03-2026 to the ONNX package used by Vernacula. See cohere_export/README.md for full documentation.
export_cohere_transcribe_to_onnx.py— exports the base Cohere encoder, mel frontend, config, tokenizer assets, and by default the KV-cache decoder pairrequirements.txt— export dependencies
Exports microsoft/VibeVoice-ASR-HF to the ONNX package used by Vernacula. See vibevoice_export/README.md for full documentation.
export_vibevoice_asr_to_onnx.py— exports the VibeVoice audio encoder, decoder graph(s), configs, and export reporttest_static_kv_parity.py— comparesdecoder_single_static.onnxagainstdecoder_single.onnxrequirements.txt— export dependencies
Exports Qwen/Qwen3-ASR-0.6B and Qwen/Qwen3-ASR-1.7B to the ONNX package we can iterate on for Vernacula. This starts from the public andrewleech/qwen3-asr-onnx export flow, trimmed down to the core export path. See qwen3asr_export/README.md for full documentation.
export_qwen3_asr_to_onnx.py— exports the Qwen3-ASR encoder, split KV-cache decoders, tokenizer assets, and config filesoptimize_qwen3_asr_graphs.py— applies ORT transformer fusions to the exported encoder and decodersprofile_qwen3_asr_pipeline.py— profiles mel, encoder, decoder prefill, and decode-step timings for an exported ONNX packagesweep_qwen3_asr_batching.py— sweeps the experimental batched encoder / decoder-prefill graphs on CUDA to map safe VRAM-dependent batch sizesrequirements.txt— export dependencies
Exports speechbrain/lang-id-voxlingua107-ecapa to a single end-to-end ONNX graph used as Vernacula's language-identification backend. See voxlingua107_export/README.md for full documentation.
export_voxlingua_to_onnx.py— exportsvoxlingua107.onnxandlang_map.jsonverify_voxlingua_parity.py— PyTorch ↔ ONNX parity check on real audio clipsrequirements.txt— export dependencies
In-progress feasibility spike for exporting AI4Bharat's IndicConformer (Hybrid CTC-RNNT, 22 Indian languages) to ONNX. Uses AI4Bharat's NeMo fork in an isolated venv. Not yet wired into the runtime — see indicconformer_export/README.md for status.
Builds subword-level KenLM ARPAs for Parakeet shallow fusion. Streams general-English and medical text corpora from HuggingFace, layers them with an upweight, tokenises with Parakeet's tokenizer, and drives lmplz. Includes a WER/CER/ROUGE-L + entity-F1 evaluation harness. See kenlm_build/README.md for the full workflow.
Standalone export of the Whisper log-mel frontend to ONNX (mel.onnx), used for parity testing and as a reference frontend. Not a full Whisper export — openai-whisper and transformers already publish complete ONNX exports.