Skip to content

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

README.md

scripts/

Developer scripts for maintaining Vernacula localisation, help files, and ONNX model exports.

The locale and help scripts require the anthropic Python package and an Anthropic API key:

pip install anthropic
export ANTHROPIC_API_KEY=sk-ant-...   # or pass --api-key on each command

Shared tools (used by every export pipeline)

To prevent format drift across the per-export directories, two scripts at the scripts/ root are the canonical entry points after a successful export run:

Typical post-export workflow:

python scripts/make_manifest.py --model-dir ~/models/<bundle> --all
python scripts/upload_to_hf.py \
    --model-dir ~/models/<bundle> \
    --repo-id christopherthompson81/<repo-name> \
    --sync-readme --create-repo

Both take --exclude GLOB... so a bundle directory that has been run from (ONNX Runtime writes *.ort / *.use-ort optimisation caches beside the graphs) can still be published with --all:

python scripts/make_manifest.py --model-dir ~/models/kokoro --all --exclude '*.ort' '*.use-ort'
python scripts/upload_to_hf.py --model-dir ~/models/kokoro \
    --repo-id christopherthompson81/kokoro-82m-onnx \
    --exclude '*.ort' '*.use-ort' --sync-readme --create-repo

The C# contract is locked at Vernacula.Avalonia/Services/ModelManagerService.cs::ParseManifestHashes — don't fork the manifest shape per-export. New shared utilities go in _export_utils/.


update_locale_keys.py

Translates new or changed UI string keys from en.json into all 24 supported language JSON files.

Before running, edit the two constants near the top of the file:

  • NEW_KEYS — the keys to add, in insertion order, with their English values
  • ANCHOR_KEY — the existing key after which the new keys will be inserted

The script only retranslates keys whose English source has changed since the last run (SQLite cache at src/Vernacula.Avalonia/Locales/locale_keys.db). Interrupted runs resume automatically.

# Translate any new/changed keys and write them into the locale JSON files
python scripts/update_locale_keys.py

# Force retranslation of all keys in NEW_KEYS, ignoring the cache
python scripts/update_locale_keys.py --force

# Pass the API key directly instead of via environment variable
python scripts/update_locale_keys.py --api-key sk-ant-...

update_help_files.py

Translates new or changed English help Markdown files (Help/en/**/*.md) into all 24 supported language directories (Help/{lang}/).

Change detection is based on SHA-256 of the English file contents (cache at src/Vernacula.Avalonia/Help/help_files.db). The prompt instructs the model to preserve YAML frontmatter structure, Markdown syntax, relative link hrefs, backtick UI labels, and technical terms — only human-readable prose is translated.

# Translate all changed files for all languages
python scripts/update_help_files.py

# Force retranslation of everything, ignoring the cache
python scripts/update_help_files.py --force

# Retranslate one specific file for all languages
python scripts/update_help_files.py --files operations/editing_transcripts.md

# Process multiple specific files (comma-separated, relative to Help/en/)
python scripts/update_help_files.py --files index.md,operations/editing_transcripts.md

# Target only specific languages
python scripts/update_help_files.py --lang de,fr

# Combine filters — one file, two languages
python scripts/update_help_files.py --files index.md --lang bg,cs

nemo_export/

Scripts for exporting NeMo model checkpoints to ONNX. See nemo_export/README.md for full documentation.

  • export_parakeet_nemo_to_onnx.py — exports a Parakeet RNNT/TDT .nemo checkpoint to the ONNX package used by Vernacula
  • export_sortformer_nemo_to_onnx.py — exports a streaming Sortformer diarization .nemo checkpoint to ONNX
  • export_silero_vad_to_onnx.py — exports Silero VAD to ONNX
  • setup_nemo_export_env.py — creates the Python export venv
  • requirements.txt — export dependencies

deepfilternet3_export/

Exports DeepFilterNet3 as three streaming ONNX models with explicit GRU hidden-state I/O for chunk-by-chunk C# inference. See deepfilternet3_export/README.md for full documentation.


diarizen_export/

Exports the DiariZen diarization pipeline: segmentation model, WeSpeaker embedding model, and LDA/PLDA transform parameters. See diarizen_export/README.md for full documentation.


cohere_export/

Exports CohereLabs/cohere-transcribe-03-2026 to the ONNX package used by Vernacula. See cohere_export/README.md for full documentation.

  • export_cohere_transcribe_to_onnx.py — exports the base Cohere encoder, mel frontend, config, tokenizer assets, and by default the KV-cache decoder pair
  • requirements.txt — export dependencies

vibevoice_export/

Exports microsoft/VibeVoice-ASR-HF to the ONNX package used by Vernacula. See vibevoice_export/README.md for full documentation.

  • export_vibevoice_asr_to_onnx.py — exports the VibeVoice audio encoder, decoder graph(s), configs, and export report
  • test_static_kv_parity.py — compares decoder_single_static.onnx against decoder_single.onnx
  • requirements.txt — export dependencies

qwen3asr_export/

Exports Qwen/Qwen3-ASR-0.6B and Qwen/Qwen3-ASR-1.7B to the ONNX package we can iterate on for Vernacula. This starts from the public andrewleech/qwen3-asr-onnx export flow, trimmed down to the core export path. See qwen3asr_export/README.md for full documentation.

  • export_qwen3_asr_to_onnx.py — exports the Qwen3-ASR encoder, split KV-cache decoders, tokenizer assets, and config files
  • optimize_qwen3_asr_graphs.py — applies ORT transformer fusions to the exported encoder and decoders
  • profile_qwen3_asr_pipeline.py — profiles mel, encoder, decoder prefill, and decode-step timings for an exported ONNX package
  • sweep_qwen3_asr_batching.py — sweeps the experimental batched encoder / decoder-prefill graphs on CUDA to map safe VRAM-dependent batch sizes
  • requirements.txt — export dependencies

voxlingua107_export/

Exports speechbrain/lang-id-voxlingua107-ecapa to a single end-to-end ONNX graph used as Vernacula's language-identification backend. See voxlingua107_export/README.md for full documentation.

  • export_voxlingua_to_onnx.py — exports voxlingua107.onnx and lang_map.json
  • verify_voxlingua_parity.py — PyTorch ↔ ONNX parity check on real audio clips
  • requirements.txt — export dependencies

indicconformer_export/

In-progress feasibility spike for exporting AI4Bharat's IndicConformer (Hybrid CTC-RNNT, 22 Indian languages) to ONNX. Uses AI4Bharat's NeMo fork in an isolated venv. Not yet wired into the runtime — see indicconformer_export/README.md for status.


kenlm_build/

Builds subword-level KenLM ARPAs for Parakeet shallow fusion. Streams general-English and medical text corpora from HuggingFace, layers them with an upweight, tokenises with Parakeet's tokenizer, and drives lmplz. Includes a WER/CER/ROUGE-L + entity-F1 evaluation harness. See kenlm_build/README.md for the full workflow.


whisper_export/

Standalone export of the Whisper log-mel frontend to ONNX (mel.onnx), used for parity testing and as a reference frontend. Not a full Whisper export — openai-whisper and transformers already publish complete ONNX exports.