You can run a useful AI assistant on your own computer with a local model runner such as Ollama and an open-weight language model. Your questions and documents then stay on your disk, the assistant works without an internet connection, and there is no per-request bill. The trade-offs are answer quality and speed. Local models are smaller than the largest hosted ones, and without a suitable graphics card replies are slow.
#Why run an assistant locally
Three reasons come up most often.
- Privacy. A hosted assistant sends what you type, and any file you attach, to the provider's servers. That may conflict with a client contract, an internal policy or simple caution about personal data. A local model processes everything on your machine, so there is nothing to hand to a third party.
- Offline use. Once the model is downloaded, it runs without a connection. That helps on a train, at a site with poor coverage, or on a machine that is kept off the internet on purpose.
- Cost control. Hosted models charge by subscription or by usage. A local model costs the hardware you already own and the electricity it draws. Heavy use, long documents and automated tasks do not add to a bill.
A fourth reason matters more over time. A local assistant can build long-term knowledge of your own files and projects without uploading them anywhere.
#What you need to run a model locally
A local setup has two parts, a program that runs the model and the model itself, plus hardware that can carry both.
A model runner. Ollama is a common choice. It downloads open-weight models, runs them on your CPU or GPU, and exposes a local API that other programs can call. Other runners exist, and most assistants that support local models can talk to at least one of them.
A model. Open-weight models come in several sizes, measured in parameters, and usually in quantized versions that trade a little quality for much lower memory use. A larger model gives better answers and needs more memory and more computing power. Start with a small or mid-sized model, see whether the answers are good enough for your tasks, and move up only if they are not.
Hardware. In general terms:
- Memory decides which models you can load at all. The model has to fit in RAM, or in the graphics card's memory if you use a GPU.
- A GPU with enough memory for the model makes replies much faster. On a CPU alone, models still run, but slowly.
- Model files are large, so plan disk space for more than one model if you want to compare them.
The model runner's documentation lists what each model needs. Check that against your machine before you download anything, and measure the speed yourself with the kind of prompt you will actually use.
#The limits compared with hosted models
Be clear about what you give up before you decide.
- Answer quality. The largest hosted models are much bigger than what a desktop computer can run. For complex reasoning, long code changes or careful writing, a local model is usually weaker.
- Speed. Without a strong GPU, a long prompt can take a noticeable time to process before the reply even starts. The length of the input often matters more than the length of the answer.
- Context. Local models tend to handle less text at once, so you cannot always paste a whole document and ask about it.
- Current knowledge. A model knows what was in its training data. Without a search tool it has no view of recent events.
- Maintenance. You update the runner and the models yourself, and you decide which models to trust.
Some tools offer a middle path: a local model by default, with an option to send a question to a hosted model when you need more capability. If you use such a mode, check exactly what leaves the machine in it.
#What makes a local assistant useful day to day
A plain chat window over a local model keeps your data private, yet it knows nothing about your work and forgets each conversation when it ends. The features that make a local assistant worth running sit around the model:
- Memory. A store of what it has learned about your projects and documents, kept on your disk, that it can recall in later sessions.
- Control over what it learns. You should be able to see and approve what it remembers, and to move or delete that memory.
- Tools with limits. Reading files or running commands is useful, and it needs permissions, a log of every action and a way to undo changes.
- A clear privacy setting. One setting that states what may leave the machine, with a safe default. Also check whether the tool sends telemetry.
When you evaluate any local assistant, open its settings and documentation and confirm each of these before you point it at private files.
#PN Brain as one open-source example
PN Brain is a free, open-source assistant that PN Scripts publishes on GitHub under the MIT license. It is one Go program with one SQLite file, and it uses a model you run in Ollama. It needs no database server, no containers and no cloud account.
- Setup. You need Ollama with a chat model and the
nomic-embed-textembedding model. The first run finds Ollama, checks what is missing and offers to install it. On Linux it opens a native window through WebKitGTK; on macOS and Windows the interface opens in the browser. - Memory. It learns from your code and documents with your approval and answers with real file paths from your disk. Anything it inferred waits for your approval, and facts it can check against the disk are saved automatically. The memory lives in a folder that can sit on an external disk and move to another computer.
- Privacy modes. In the default private mode nothing leaves the machine. Research mode sends search queries only, and open mode may send what you type to a third-party model. In every mode, what it has learned about you stays on your computer. There is no telemetry, and an unknown privacy value is treated as private.
- Tools. It can read files, list folders, run commands, search the web and control a smart home through Home Assistant. Every action is recorded, a changed file can be restored, and you choose how often it asks for permission.
- Voice. Optional voice conversation uses whisper.cpp to listen and local speech output to answer. Spoken turns cannot take actions, so a misheard sentence cannot change a file.
The product page is also direct about speed. On a CPU without a GPU replies are slow, and the page publishes measured figures so you know what to expect before you install it.
#How PN Scripts can help
If you want to try a private assistant that remembers your projects, the PN Brain product page lists the requirements, explains the privacy modes and links to the code on GitHub. PN Scripts can also adapt it for a team or a project as paid work.
Comments
Comments
Be the first to leave a comment.
Leave a comment