Skip to content

Local AI models

Run summaries, tasks, and chat on your own computer with Kimiko’s built-in AI. Compare Gemma and Qwen models, or connect Ollama or LM Studio.

Last verified For Kimiko 0.5.9

Open Settings → AI → Built-in AI. Download a model, then select the one you want your AI tasks to use.

The recommended onboarding choice is Gemma 4 E2B. It is a useful starting point because Kimiko can set up its local AI workflow for you.

Model Approximate download When to consider it
Gemma 4 E2B 3.2 GB The recommended onboarding setup
Qwen 3.5 2B 1.2 GB When you want a smaller download
Qwen 3.5 4B 2.6 GB To compare another local model on your notes
Gemma 4 E4B 4.9 GB To explore a larger Gemma option
Qwen 3.5 9B 5.4 GB To explore a larger Qwen option
Gemma 4 12B 6.7 GB A larger download to test if your computer has the resources

These are the approximate downloads shown in Kimiko’s catalog. They do not describe peak memory use or guarantee speed or answer quality. Start with one model and test it with your typical meetings before downloading more.

The AI tasks section controls model choices for different jobs. Check those choices after changing providers, especially if you want all processing to remain local.

Dictation formatting can follow its AI task choice or use an override in Settings → Dictation. Plain dictation transcription is a separate speech-model choice.

The built-in setup uses EmbeddingGemma 300M, an additional download of approximately 329 MB, to help find relevant passages for chat. It is a search model, not the model that writes the answer.

Check Settings → AI → Chat search index if recently added meetings are missing from answers. Indexing and answer generation have separate roles.

If you already run Ollama or LM Studio, Kimiko provides connection presets under AI providers. Connect the server you actually run, test the connection, and choose its model for the intended tasks.

For the fewest setup steps, use Kimiko’s built-in AI first. You can change this later.

Close memory-heavy apps, try a smaller model, and compare a short meeting. A model’s file size is only part of its resource needs. See slow processing.

If your laptop has shared memory (UMA), AI and graphics draw from the same RAM as your other apps. Memory use on laptops explains why usage can be higher and what you can adjust.