Home›Resources›Guides
Local AI

Local AI on your Mac

Keep every request, page and file on your Mac. Here's what runs, how Linda picks it, and what to expect.

Updated 30 September 2026 · Linda 0.2.0

Why run it locally

With an assistant on your Mac, the task, the pages Linda reads and the files it opens never leave the machine. There's no key, no account and no bill, and it keeps working on a plane. The trade-off is speed and skill: small models answer well and handle files, while long multi-site tasks go better with larger models or an online one.

How Linda picks an assistant

Linda reads your chip (M1 to M5, base to Ultra) and memory. An assistant fits when its download plus about 1.7 GB stays within the memory macOS lets the GPU use. Linda then estimates how fast it will write on your chip's memory bandwidth and recommends the strongest one that stays quick (about 12 words a second or more). A 4B assistant writes about 60 tokens a second on an M2 Pro.

Settings › Assistant: the assistants Linda runs on this Mac, the recommended one first

Settings › Assistant shows Linda's recommended assistants first, then More assistants: every other open assistant that can use Linda's tools, searchable. Each row says whether it runs well, “Leaves little memory for your apps”, or needs a bigger Mac. Assistants are MLX builds from Hugging Face and run on MLX, Apple's machine-learning framework for Apple silicon. The first time you download one, Linda installs its MLX engine privately (about 0.5 GB, once) and keeps it up to date; nothing touches your own Python or Homebrew. The conversation length is set from your memory and can be changed per assistant.

Already use Ollama or LM Studio?

Choose Advanced in Settings › Assistant. Linda finds the app on your Mac, starts it when a task needs it, and talks to it over the OpenAI-compatible API:

AppAddressExample model
Ollama127.0.0.1:11434qwen3.5:9b
LM Studio127.0.0.1:1234qwen/qwen3.5-9b
mlx-lm127.0.0.1:8080mlx-community/Qwen3.5-9B-4bit

Only local addresses are accepted in this mode, so it can't be pointed at a server somewhere else by mistake. For web lookups, the assistant uses Linda's browser.

System One: quick decisions

Most browser steps are a pick among the page's own actions. A decision model scores those options in one pass instead of writing an answer, so it decides in a fraction of a second. Laya, an open typed-decision model (0.85 GB), downloads by itself and runs inside Linda in about 50 ms a decision; Ollama 0.35's System One endpoint with Nimble works too. When it's unsure (under 60%), or the step needs typing, your assistant takes over. Every pick still passes Linda's approval rules. How it works →

Tips

  • Close other heavy apps before the first download; the recommended assistant assumes most of your memory is free.
  • Thinking is off by default for speed. Turn it on in Settings if you prefer careful answers.
  • Use a project for related tasks, so files carry over and the assistant has less to re-read.
  • Mix and match: keep local for private work and switch to an online assistant for one hard research task.

Try Linda on your Mac

Free for any Mac with Apple silicon and macOS 26. No account.

Download for Mac