Why run it locally
With an assistant on your Mac, the task, the pages Linda reads and the files it opens never leave the machine. There's no key, no account and no bill, and it keeps working on a plane. The trade-off is speed and skill: small models answer well and handle files, while long multi-site tasks go better with larger models or an online one.
How Linda picks an assistant
Linda reads your chip (M1 to M5, base to Ultra) and memory. An assistant fits when its download plus about 1.7 GB stays within the memory macOS lets the GPU use. Linda then estimates how fast it will write on your chip's memory bandwidth and recommends the strongest one that stays quick (about 12 words a second or more). A 4B assistant writes about 60 tokens a second on an M2 Pro.

Settings › Assistant shows Linda's recommended assistants first, then More assistants: every other open assistant that can use Linda's tools, searchable. Each row says whether it runs well, “Leaves little memory for your apps”, or needs a bigger Mac. Assistants are MLX builds from Hugging Face and run on MLX, Apple's machine-learning framework for Apple silicon. The first time you download one, Linda installs its MLX engine privately (about 0.5 GB, once) and keeps it up to date; nothing touches your own Python or Homebrew. The conversation length is set from your memory and can be changed per assistant.
Already use Ollama or LM Studio?
Choose Advanced in Settings › Assistant. Linda finds the app on your Mac, starts it when a task needs it, and talks to it over the OpenAI-compatible API:
| App | Address | Example model |
|---|---|---|
| Ollama | 127.0.0.1:11434 | qwen3.5:9b |
| LM Studio | 127.0.0.1:1234 | qwen/qwen3.5-9b |
| mlx-lm | 127.0.0.1:8080 | mlx-community/Qwen3.5-9B-4bit |
Only local addresses are accepted in this mode, so it can't be pointed at a server somewhere else by mistake. For web lookups, the assistant uses Linda's browser.
System One: quick decisions
Most browser steps are a pick among the page's own actions. A decision model scores those options in one pass instead of writing an answer, so it decides in a fraction of a second. Laya, an open typed-decision model (0.85 GB), downloads by itself and runs inside Linda in about 50 ms a decision; Ollama 0.35's System One endpoint with Nimble works too. When it's unsure (under 60%), or the step needs typing, your assistant takes over. Every pick still passes Linda's approval rules. How it works →
Tips
- Close other heavy apps before the first download; the recommended assistant assumes most of your memory is free.
- Thinking is off by default for speed. Turn it on in Settings if you prefer careful answers.
- Use a project for related tasks, so files carry over and the assistant has less to re-read.
- Mix and match: keep local for private work and switch to an online assistant for one hard research task.