Skip to content
PN Scripts

Open source

Running an AI assistant on your own computer: what you need and what to expect

  • PN Scripts Team
  • 6 min read

You can run a useful AI assistant on your own computer with a local model runner such as Ollama and an open-weight language model. Your questions and documents then stay on your disk, the assistant works without an internet connection, and there is no per-request bill. The trade-offs are answer quality and speed. Local models are smaller than the largest hosted ones, and without a suitable graphics card replies are slow.

#Why run an assistant locally

Three reasons come up most often.

  • Privacy. A hosted assistant sends what you type, and any file you attach, to the provider's servers. That may conflict with a client contract, an internal policy or simple caution about personal data. A local model processes everything on your machine, so there is nothing to hand to a third party.
  • Offline use. Once the model is downloaded, it runs without a connection. That helps on a train, at a site with poor coverage, or on a machine that is kept off the internet on purpose.
  • Cost control. Hosted models charge by subscription or by usage. A local model costs the hardware you already own and the electricity it draws. Heavy use, long documents and automated tasks do not add to a bill.

A fourth reason matters more over time. A local assistant can build long-term knowledge of your own files and projects without uploading them anywhere.

#What you need to run a model locally

A local setup has two parts, a program that runs the model and the model itself, plus hardware that can carry both.

A model runner. Ollama is a common choice. It downloads open-weight models, runs them on your CPU or GPU, and exposes a local API that other programs can call. Other runners exist, and most assistants that support local models can talk to at least one of them.

A model. Open-weight models come in several sizes, measured in parameters, and usually in quantized versions that trade a little quality for much lower memory use. A larger model gives better answers and needs more memory and more computing power. Start with a small or mid-sized model, see whether the answers are good enough for your tasks, and move up only if they are not.

Hardware. In general terms:

  • Memory decides which models you can load at all. The model has to fit in RAM, or in the graphics card's memory if you use a GPU.
  • A GPU with enough memory for the model makes replies much faster. On a CPU alone, models still run, but slowly.
  • Model files are large, so plan disk space for more than one model if you want to compare them.

The model runner's documentation lists what each model needs. Check that against your machine before you download anything, and measure the speed yourself with the kind of prompt you will actually use.

#The limits compared with hosted models

Be clear about what you give up before you decide.

  • Answer quality. The largest hosted models are much bigger than what a desktop computer can run. For complex reasoning, long code changes or careful writing, a local model is usually weaker.
  • Speed. Without a strong GPU, a long prompt can take a noticeable time to process before the reply even starts. The length of the input often matters more than the length of the answer.
  • Context. Local models tend to handle less text at once, so you cannot always paste a whole document and ask about it.
  • Current knowledge. A model knows what was in its training data. Without a search tool it has no view of recent events.
  • Maintenance. You update the runner and the models yourself, and you decide which models to trust.

Some tools offer a middle path: a local model by default, with an option to send a question to a hosted model when you need more capability. If you use such a mode, check exactly what leaves the machine in it.

#What makes a local assistant useful day to day

A plain chat window over a local model keeps your data private, yet it knows nothing about your work and forgets each conversation when it ends. The features that make a local assistant worth running sit around the model:

  • Memory. A store of what it has learned about your projects and documents, kept on your disk, that it can recall in later sessions.
  • Control over what it learns. You should be able to see and approve what it remembers, and to move or delete that memory.
  • Tools with limits. Reading files or running commands is useful, and it needs permissions, a log of every action and a way to undo changes.
  • A clear privacy setting. One setting that states what may leave the machine, with a safe default. Also check whether the tool sends telemetry.

When you evaluate any local assistant, open its settings and documentation and confirm each of these before you point it at private files.

#PN Brain as one open-source example

PN Brain is a free, open-source assistant that PN Scripts publishes on GitHub under the MIT license. It is one Go program with one SQLite file, and it uses a model you run in Ollama. It needs no database server, no containers and no cloud account.

  • Setup. You need Ollama with a chat model and the nomic-embed-text embedding model. The first run finds Ollama, checks what is missing and offers to install it. On Linux it opens a native window through WebKitGTK; on macOS and Windows the interface opens in the browser.
  • Memory. It learns from your code and documents with your approval and answers with real file paths from your disk. Anything it inferred waits for your approval, and facts it can check against the disk are saved automatically. The memory lives in a folder that can sit on an external disk and move to another computer.
  • Privacy modes. In the default private mode nothing leaves the machine. Research mode sends search queries only, and open mode may send what you type to a third-party model. In every mode, what it has learned about you stays on your computer. There is no telemetry, and an unknown privacy value is treated as private.
  • Tools. It can read files, list folders, run commands, search the web and control a smart home through Home Assistant. Every action is recorded, a changed file can be restored, and you choose how often it asks for permission.
  • Voice. Optional voice conversation uses whisper.cpp to listen and local speech output to answer. Spoken turns cannot take actions, so a misheard sentence cannot change a file.

The product page is also direct about speed. On a CPU without a GPU replies are slow, and the page publishes measured figures so you know what to expect before you install it.

#How PN Scripts can help

If you want to try a private assistant that remembers your projects, the PN Brain product page lists the requirements, explains the privacy modes and links to the code on GitHub. PN Scripts can also adapt it for a team or a project as paid work.

Keep reading

Keep reading

Comments

Comments

Be the first to leave a comment.

Leave a comment

Next step

pnscripts.com/contact

Talk to the team

Ask about an article, or tell us about a project you want built. We reply within one business day.

Write to us

The PN Scripts family

Other PN Scripts sites

Hosting, games and the blog each have their own site, run by the same company.

  • pnscripts.com

    PN Scripts

    Software engineering

    Custom web, mobile, API and game development, plus our open-source products and plugins.

  • games.pnscripts.com

    Games

    Games and game servers

    The home for PN Scripts games and game servers. The catalog is empty for now and fills up as titles and servers go live.

  • hosting.pnscripts.com

    Hosting

    Hosting and infrastructure

    Shared hosting, KVM VPS, dedicated servers and domains, from the same company that builds your project.

  • blog.pnscripts.com

    Blog

    Articles and field notes

    Plain articles on hosting, servers, domains and security, written by the people who work with them.

    You are here