Memory use on laptops
Why Kimiko can use more RAM on laptops with shared memory (UMA), even with cloud processing, and the settings that lighten the load.
Last verified For Kimiko 0.6.7
Kimiko can use several gigabytes of memory while transcribing a meeting or writing a summary locally. On a laptop with shared memory, that work draws from the same memory your other apps need. The amount depends on your models, enabled features, and computer.
For a quick hardware checklist, see system requirements. A 16 GB Apple silicon Mac or Windows PC with dedicated graphics is a practical starting point for local speech and AI. Windows shared-memory laptops may benefit from more room.
What shared memory means
Section titled “What shared memory means”Many Windows laptops have built-in graphics that share the computer’s main memory (RAM). The main processor and graphics processor use one pool of memory. This is called unified memory architecture, or UMA. AMD explains this shared-memory design.
For example, on a laptop with 16 GB of shared memory, Windows, your apps, and graphics all draw from that 16 GB. A computer with a separate graphics card can keep some AI data in the card’s own memory instead. The laptop can therefore use more system RAM for the same task.
Apple silicon Macs also use shared memory, but their hardware and software handle local processing differently. Shared memory alone does not mean high usage or poor performance; Kimiko has been used successfully on a 16 GB M5 MacBook Air.
Sharing memory does not guarantee that every part of the software shares a single copy of its data. AI models can also need extra copies and working space while they run. A model’s download size is only its size on disk.
Why usage rises during a meeting
Section titled “Why usage rises during a meeting”Kimiko may be doing several jobs: turning speech into text, separating speakers, preparing notes for search, and writing a summary. Each needs memory. A local summary can temporarily raise usage while the speech model is still running.
These jobs also use processing power, so other demanding apps can slow them down. A high memory reading alone does not mean something is broken; check whether your computer stays responsive and the transcript keeps up.
Why cloud processing still uses memory
Section titled “Why cloud processing still uses memory”Cloud speech and cloud AI are separate choices. Moving one to the cloud leaves the other unchanged.
Even with both selected, Kimiko still handles audio and detects speech on your computer. Speaker labels and voice recognition run locally when enabled. Chat search may also use a local model, and enabled dictation has its own speech model. Cloud processing can reduce the load, but it does not remove every local job. See cloud providers.
How Kimiko reduces the load
Section titled “How Kimiko reduces the load”- Since 0.6.5: on Windows, speech detection runs on the main processor. On computers with only shared-memory graphics, speaker labeling uses it too, avoiding extra graphics-related overhead for those jobs. Transcription and local AI can still use graphics acceleration.
- Since 0.6.6: Kimiko unloads the waiting dictation model before a meeting starts. On Windows computers with only shared-memory graphics, it also changes how built-in AI loads model data so Windows can reclaim some memory more easily.
- Since 0.6.7: dictation’s speech and built-in formatting models use a shared 5-minute idle window by default. Change Keep models loaded in Settings → Dictation to free memory sooner or keep models ready longer. A model also used by chat may stay loaded longer.
These changes happen automatically. They do not remove the memory needed by active models, and there is no single memory target that applies to every laptop.
What you can try
Section titled “What you can try”- Update Kimiko to include the changes above.
- Close demanding apps and compare a short meeting with a smaller local AI model.
- Use fewer optional features. In Settings → Transcription, turn off Label remote speakers if you do not need separate speaker labels, or Live summaries if you can wait and summarize after recording.
- Consider cloud processing for speech, AI, or both if it suits your privacy preferences. Review the separate choices under AI tasks too.
If memory keeps growing across meetings, or the computer stays slow after processing finishes, contact support with your Kimiko version, laptop model, selected models, and enabled features.
Reading Windows Task Manager
Section titled “Reading Windows Task Manager”Use Performance → Memory to check overall RAM use. Kimiko’s main app, AI engine, and window display can appear as separate processes. Shared GPU memory already comes from system RAM; adding it to total RAM usage would count it twice.
A quiet 3D graph does not prove the graphics processor is idle. Under Performance → GPU, click a graph’s name and select a Compute graph, such as Compute 0, if available. AI work may appear there. Microsoft explains Task Manager’s GPU views.