Skip to content

Models

Herga uses two local models, both running on your Mac’s GPU through Apple’s MLX framework:

  • Whisper turns your speech into text.
  • Qwen3 cleans that text up.

Choose them in Settings › Transcription & refinement, and manage downloads in the Models tab.

Model Size on disk Notes
Base 290 MB Fastest, least accurate
Small 970 MB A good fallback for slow connections or small disks
Medium 3.1 GB
Large v3 3.1 GB Most accurate, slowest
Turbo 1.6 GB Default. Close to Large in quality and much faster

Turbo is the right choice for almost everyone. Whisper stays loaded while Herga is running, so your first dictation is as fast as the rest.

Language: Auto-detect, English, Spanish, French, German, Japanese, Chinese or Hindi. Picking your language is a little faster and more reliable than auto-detect.

Model Size on disk Notes
0.6B 400 MB Default. Fast, and good at everyday cleanup
1.7B 1.1 GB Better with long, rambling dictations
4B 2.5 GB Best at following your style and at Command Mode rewrites

Larger models give better results but take longer and use more memory.

To skip cleanup and get Whisper’s text as-is, turn off Refine transcripts automatically. You can still refine any capture later.

From the Models tab you can download, cancel, unload and delete models, and Show in Finder. Problems with a model are flagged here too.

Models live in the Hugging Face cache on your Mac. To store them somewhere else, like an external drive, click Move… in the Models tab. Herga stops its local server during the move, shows the progress, and starts again when it’s done. Reset moves them back to the default location.

You can also set the HERGA_MODELS_DIR environment variable to choose the folder.