Skip to content

Choose your model setup

Start with Kimiko’s recommended local speech and AI models, see what to change for your needs, and keep all processing on your computer if you prefer.

Last verified For Kimiko 0.5.9

For your first meeting, use the speech model offered in onboarding and the recommended built-in AI option. This gives you a local starting point without connecting a cloud account.

These recommendations follow Kimiko’s built-in defaults. They are not a hardware benchmark or a ranking of every available model.

Job Starting choice Where to change it
Meeting transcription Nemotron 3.5 Streaming Settings → Transcription
Summaries, tasks, and chat Built-in Gemma 4 E2B Settings → AI
Dictation speech Fastest, using Parakeet v3 Settings → Dictation
Dictation cleanup Clean for a simple starting point Settings → Dictation → Formatting

Download the required models before your first call. Model availability and language support are shown in the app’s pickers.

Speech models turn audio into text. They affect how your meetings and dictation are transcribed.

AI models work with text. They help with summaries, task suggestions, questions, and optional dictation formatting.

You can run both locally, choose cloud processing for one, or configure both to use cloud services. Choosing cloud AI does not automatically change your speech model.

If you want to… Try this
Keep processing on your computer Use local speech and built-in AI; review local-only mode after downloads finish
Reduce the size of the AI download Consider Qwen 3.5 2B; compare its results on your own notes
Explore another local AI model Compare a Gemma or Qwen option in local AI
Improve recognition for your language Check language support and try a different speech model
Use a cloud model you already have access to Connect it in cloud providers and choose which AI tasks use it

Check the system requirements before choosing a fully local setup. A 16 GB Apple silicon Mac or Windows PC with dedicated graphics is a practical starting point. Larger models and heavy multitasking benefit from more memory; the guide explains the tradeoffs.

Try a short, representative meeting in your usual language. Check the transcript’s names and numbers, the summary’s decisions, and how long processing takes while your other apps are open.

Download size is not the amount of memory a model needs while running. Larger models can require more memory and may take longer. If your computer struggles, try a smaller option and compare the result.

On a laptop with shared memory (UMA), graphics use the same RAM as your apps. Read memory use on laptops for practical adjustments and an explanation of why cloud processing still uses some local resources.

There is no need to download every model. Keep a setup that works well for your actual conversations.