# Ollama using too much memory on a Mac? How to see what is loaded, and unload it

Canonical page: https://vitalsmac.com/ollama-memory-usage-mac

By Atta Ur Rehman Shah, maker of Vitals. Published October 3, 2026.

**Short answer:** Ollama keeps a model in memory for five minutes after its last answer, so the next question is answered quickly. That is why it still holds several gigabytes when you have stopped using it. Run `ollama ps` to see which models are loaded and when each will be unloaded, and `ollama stop` followed by a model’s name to unload it now. How much memory a model takes depends on its size and on the context length you run it with. Vitals, a system monitor for the Mac, shows Ollama’s memory, the graphics chip’s load and the Mac’s memory pressure in the menu bar.

## How to see and free Ollama’s memory

1. In Terminal, run `ollama ps`. Each row is a model in memory, with its size, whether it runs on the graphics chip, and how long until it is unloaded.
2. Unload one now with `ollama stop` followed by its name, for example `ollama stop llama3.2`. The memory comes back within a second or two.
3. Open Activity Monitor, click Memory, and look at Memory Pressure at the bottom. Green means the model fits. Yellow or red while a model is loaded means it is too large for what else is open.
4. To have models unload sooner, set how long they are kept. In Terminal, run `launchctl setenv OLLAMA_KEEP_ALIVE 1m`, then quit Ollama from its menu bar icon and open it again.
5. If the Mac is short of memory whenever a model is loaded, use a smaller model or a smaller version of the same one. `ollama list` shows the ones you have and their sizes.

A keep-alive of `0` unloads a model as soon as it has answered; `-1` keeps it loaded until you stop it. Apps that talk to Ollama can set their own value with each request, which is why a model sometimes stays longer than your setting.

Running a model on your own Mac means holding all of it in memory at once. A Mac with Apple silicon shares one pool of memory between the processor and the graphics chip, so a model that would need a dedicated graphics card elsewhere fits here, but it comes out of the same memory your apps use. An 8 GB model on a 16 GB Mac leaves 8 GB for everything else. Ollama is not leaking when it holds that much; it is holding the model. The table shows what decides the number.

## What decides Ollama’s memory

| What you see | Why | What to do |
|---|---|---|
| Several GB held after you stop chatting | The model stays loaded for five minutes | `ollama stop`, or a shorter keep-alive |
| More memory than the model’s file size | The context you run it with needs memory too | Use a smaller context length |
| Two models in memory at once | A second model was asked for before the first was unloaded | `ollama ps`, then stop the one you are not using |
| Memory pressure yellow or red, the Mac slow | The model is too large for the memory that is free | A smaller model, or close other apps |
| Slow answers and a busy CPU | Part of the model did not fit on the graphics chip | `ollama ps` shows the split; use a smaller model |

## Why Ollama keeps models loaded

Loading a model means reading gigabytes from disk into memory, which takes seconds to a minute. Doing that for every question would make every answer slow, so Ollama keeps the model for five minutes after its last use and unloads it after that. Each new question resets the clock. If an app on your Mac asks the model something every few minutes, an editor’s autocomplete for instance, the model never unloads.

The setting is called keep-alive. `OLLAMA_KEEP_ALIVE` changes the default for everything; a value of 0 gives the memory back at once at the cost of a slow first answer every time.

## How much memory a model needs

As a rule the memory is the size of the model’s file plus room for the conversation. A model with 8 billion parameters at the usual compression is about 5 GB; one with 30 billion is about 20 GB. The size on disk that `ollama list` shows is a good first estimate.

Context length is the part people miss. It is how much text the model can hold in mind at once, and memory for it is set aside when the model loads, whether or not you fill it. A long context on a small model can use more than the model itself. If `ollama ps` shows a size well above the file’s, the context is why.

## Does the model fit your Mac?

The honest test is [memory pressure](https://vitalsmac.com/glossary/memory-pressure) with the model loaded and your usual apps open. Green: it fits. Yellow: macOS is compressing memory to make room, and things slow down a little. Red: it is [swapping to disk](https://vitalsmac.com/glossary/swap-used), and both the model and the Mac become slow. A model that runs partly on the processor because it did not fit on the graphics chip is also much slower; `ollama ps` says how it was split.

A smaller version of the same model usually costs little in quality and a lot less in memory. On a Mac with 16 GB, models up to about 8 GB leave room to work. See [unified memory](https://vitalsmac.com/glossary/unified-memory) for how the Mac shares it.

## LM Studio and other apps

The same applies to any app that runs models locally. LM Studio keeps a model loaded until you eject it or its idle time runs out, and shows the loaded models at the top of its window. Editors and chat apps that run a model for you do the same behind the scenes. The memory is the model, whichever app holds it.

## What Vitals adds

Vitals shows what a loaded model costs the Mac, where you can see it while you work.

- Ollama’s memory as one row, its helper processes counted in
- Memory pressure and free memory in the menu bar, so you see at once whether a model fits
- The graphics chip’s load in the menu bar and its own tab, with the apps using it
- A notification when an app’s memory keeps growing
- Which app used the most energy, hour by hour: running a model is among the heaviest things a Mac does on battery

Good to know:

- Vitals shows the memory Ollama holds, not which model is loaded. `ollama ps` shows that.
- Vitals does not unload models. `ollama stop` does.

Vitals is $9 during the launch, $29 after it, once: [vitalsmac.com](https://vitalsmac.com/).

## Questions people ask

### Why is Ollama using so much memory?

It is holding a model in memory. Models stay loaded for five minutes after their last answer. Run `ollama ps` to see what is loaded, and `ollama stop` with the model’s name to unload it.

### How do I unload a model in Ollama?

Run `ollama stop` followed by the model’s name. To have models unload by themselves sooner, set `OLLAMA_KEEP_ALIVE` to a shorter time.

### How do I change how long Ollama keeps a model loaded on a Mac?

Run `launchctl setenv OLLAMA_KEEP_ALIVE 1m` in Terminal, with the time you want, then quit and reopen Ollama. Use 0 to unload at once, or -1 to keep models loaded.

### How much memory do I need to run a model with Ollama?

About the size of the model’s file plus room for the context. On a 16 GB Mac, models up to about 8 GB leave room for your other apps.

### Why does Ollama use more memory than the model’s size?

Memory for the context length is set aside when the model loads. A long context can add several gigabytes. Run the model with a smaller context if you do not need it.

### Does Ollama use memory when I am not using it?

Only for five minutes after the last answer, unless an app keeps asking it something or the keep-alive time has been made longer. After that the model is unloaded and Ollama uses very little.

### Is there an app that shows whether a local model fits my Mac’s memory?

Vitals shows memory pressure, free memory and the graphics chip’s load in the menu bar, so you can see what a loaded model costs while you work.

In the glossary: [Unified memory](https://vitalsmac.com/glossary/unified-memory), [Memory pressure](https://vitalsmac.com/glossary/memory-pressure), [Swap Used](https://vitalsmac.com/glossary/swap-used), [% GPU](https://vitalsmac.com/glossary/gpu-percent), [App Memory](https://vitalsmac.com/glossary/app-memory).

More guides: [Why is my Mac so slow](https://vitalsmac.com/why-is-my-mac-so-slow), [What’s draining my Mac’s battery](https://vitalsmac.com/mac-battery-draining-fast), [Which app is using my CPU or memory](https://vitalsmac.com/check-cpu-memory-usage-mac), [Why won’t my Mac sleep](https://vitalsmac.com/mac-wont-sleep), [Chrome using too much memory](https://vitalsmac.com/chrome-using-too-much-memory-mac), [Why are my Mac’s fans so loud](https://vitalsmac.com/mac-fan-loud), [How to control fan speed on a Mac](https://vitalsmac.com/mac-fan-control), [How to free up RAM on a Mac](https://vitalsmac.com/free-up-ram-mac), [Docker using too much memory](https://vitalsmac.com/docker-using-too-much-memory-mac), [How to stop Docker containers](https://vitalsmac.com/stop-docker-containers-mac), [Volume mixer for Mac](https://vitalsmac.com/mac-volume-mixer), [Cursor using high CPU](https://vitalsmac.com/cursor-high-cpu-mac), [Claude Code using too much memory](https://vitalsmac.com/claude-code-memory-usage-mac), [Claude app using high CPU](https://vitalsmac.com/claude-desktop-high-cpu-mac), [Codex using high CPU](https://vitalsmac.com/codex-high-cpu-mac), [Activity Monitor alternatives](https://vitalsmac.com/best-activity-monitor-alternatives), [iStat Menus alternatives](https://vitalsmac.com/best-istat-menus-alternatives), [Stats alternatives](https://vitalsmac.com/best-stats-alternatives), [Task Manager for Mac](https://vitalsmac.com/task-manager-for-mac), [Best system monitors for Mac](https://vitalsmac.com/best-system-monitor-for-mac), [Best menu bar monitors for Mac](https://vitalsmac.com/best-menu-bar-system-monitor-for-mac), [Best volume mixers for Mac](https://vitalsmac.com/best-volume-mixer-for-mac).
