Speech models
Compare Kimiko’s local speech models, such as Nemotron, Parakeet, and Whisper, check language support, and choose models for meetings and dictation.
Last verified For Kimiko 0.5.9
Meeting transcription
Section titled “Meeting transcription”Open Settings → Transcription → Speech models to see local and cloud options. Download a local model before selecting it for a recording.
The onboarding default is Nemotron 3.5 Streaming. Start there, then compare another model if your language or recordings need a different fit.
| Local option | Approximate download | How to approach it |
|---|---|---|
| Nemotron 3.5 Streaming | 716 MB | The onboarding default for live transcription |
| Parakeet 0.6B v3 | 705 MB | An alternative local speech model; also used by dictation’s Fastest option |
| Whisper Small | 190 MB | A smaller download to try when disk space is limited |
| Whisper Medium | 540 MB | Another Whisper option to compare on your recordings |
| Whisper Large v3 Turbo | 575 MB | A further Whisper option; test its results and responsiveness |
This is a starting selection, not the complete catalog. Download sizes are approximate model-file sizes, not RAM requirements. Some engines show text in chunks instead of continuously updating partial words.
Match the language
Section titled “Match the language”Check a model’s language labels before choosing it. The choices differ: for example, Parakeet Unified (EN) is English-only, while Parakeet 0.6B v3 is a different model with multiple supported languages.
Set Meeting language in Transcription settings. If you work in more than one language, test a short sample containing that mix before relying on it for a full meeting.
Dictation is independent
Section titled “Dictation is independent”The dictation speech model and language live in Settings → Dictation. They do not follow the meeting model automatically.
Start with Fastest if it supports your language. See the dictation guide for formatting and shortcut settings.
Cloud speech
Section titled “Cloud speech”Kimiko also offers Deepgram and ElevenLabs speech options. Cloud transcription sends audio to your chosen provider and requires its connection details and network access.
Configure cloud speech separately from cloud AI. Follow the Deepgram and ElevenLabs setup guide to create an API key and connect your provider.
Compare before switching
Section titled “Compare before switching”Use clear audio and the same language when comparing models. Check words, names, speaker context, and responsiveness. Changing the model does not correct existing transcripts automatically.