Last autumn a small engineering firm asked me to put a document assistant on their test lab network. The lab network is deliberately not connected to anything: no internet, no route out, and the only way data gets in is on a drive that somebody signs for. They wanted engineers to be able to ask questions about their own test procedures and equipment manuals without walking back to the office.
I built the box at home, loaded the models, tested it with my Wi-Fi switched off, and drove it over feeling pleased with myself. It answered questions perfectly on the first morning. Then an engineer uploaded a PDF manual, and the chat window sat there with a spinning indicator for ninety seconds before failing with an error that mentioned a connection timeout to a website I had never heard of.
The short answer to the question in the title is yes: a local language model can run with no internet connection at all, forever, and with no account anywhere. The model is just a file, and the software that runs it does not need permission from anyone. But “the model runs offline” and “the whole setup works offline” are two different claims, and the gap between them is where my first morning went.
What genuinely needs nothing
Start with the reassuring part. Once the pieces are on the disk, the core of local inference has no reason to talk to the network:
- The model weights are a file, usually GGUF for llama.cpp-based tools. Once you have it, nobody can revoke it and nothing phones home to validate it.
- The runtime, whether llama.cpp, Ollama or LM Studio, loads that file into memory and does arithmetic on it. No licence server, no login.
- No account is required to run open-weight models. Some models are “gated” on Hugging Face, meaning you need an account and must accept the licence to download them. That is a download-time step. Running the file afterwards involves nobody. The licence terms still apply to how you use it; they just are not enforced by a server.
The purest version of an offline setup is llama.cpp’s server binary and a single GGUF file. Copy both onto a machine, run one command, and you have a chat interface and an OpenAI-compatible API on the local network. I keep that combination on a USB SSD as a known-good fallback. It has never tried to reach the internet because it has nothing to reach for.
What quietly assumes a connection
The trouble is everything wrapped around the model. Modern tooling is written by people who are always online, and many components fetch something on first use rather than at install time. These are the ones that caught me or clients I have worked with since:
Model downloads that are really registry pulls
ollama pull fetches from Ollama’s registry. On an air-gapped machine there is no registry, so you need a different path in (more on that below). LM Studio’s model browser is likewise an online catalogue; the app itself runs fine offline once models are on disk.
The embedding model nobody mentioned
This was my PDF failure. When you ask a chat front end to answer questions about documents, it does not feed the whole PDF to the language model. It splits the document into chunks and uses a second, much smaller model, an embedding model, to find the chunks relevant to each question. Open WebUI, which I was using, ships with a default embedding model that it downloads from Hugging Face the first time it is needed. At home, with my Wi-Fi off, I had never uploaded a document, so it had never tried.
On the same path there was a second fetch: a tokenizer data file used when splitting text, downloaded on first use from a cloud storage address. Neither is a secret or a bug. Both are exactly the kind of dependency that works invisibly online and fails confusingly offline.

Containers, packages and fonts
If your setup runs in Docker, the images have to get onto the machine somehow; docker pull will not work. Python-based tools will try to install packages from PyPI. Some web interfaces load fonts or scripts from public content delivery networks, which does not break anything important but produces broken-looking pages and a string of errors in the browser console.
Update checks and telemetry
Desktop apps check for updates. Offline, that usually fails silently, which is fine, but occasionally a check that times out slowly adds a delay to startup. A few tools send anonymous usage statistics; offline, those simply fail.
How I build an offline box now
After that first morning I changed my process. The key change is simple: stage everything on a connected machine, then test on a machine that is genuinely cut off before it goes anywhere near the real site. Switching off Wi-Fi on the build machine is not a real test, because I never exercised every feature while it was off.
1. Collect the files on a connected machine
- Model files. Download the GGUF files directly from Hugging Face rather than through a tool’s registry. Note the SHA-256 checksum shown on each file’s page.
- If you use Ollama, you have two options. You can pull models on a connected machine and copy the entire models directory (on Linux, the
blobsandmanifestsfolders under the models path) to the same location on the offline machine. Or you can import a GGUF directly with a small Modelfile containing aFROMline pointing at the file, then runollama create. I prefer the second, because it means the offline machine only ever sees files I downloaded and checked myself. - The embedding model and any other helper models your front end uses. Find out what they are by reading the documentation or, better, by running the stack online once while watching the logs for downloads.
- Software. Installers for the runtime;
docker savefor any container images, which you load on the other side withdocker load; and for Python tools,pip downloadto collect packages you can install offline with--no-index.
2. Verify what you carry in
Before anything goes onto the offline machine, I compute the SHA-256 of each model file and compare it with the value from the download page. A multi-gigabyte file that got corrupted on a cheap USB stick does not always fail to load. Sometimes it loads and produces subtly wrong output, which is a miserable thing to diagnose on a machine where you cannot search for the error message.
3. Tell the software it is offline
Many tools can be told to stop trying. Libraries built on Hugging Face’s tooling respect the HF_HUB_OFFLINE=1 environment variable, which makes them use only files already in the local cache and fail fast instead of waiting on a timeout. Recent versions of Open WebUI have an offline mode setting and an option to stop it trying to update its embedding model automatically. Point the front end at a local copy of the embedding model rather than a download name.

4. Test on a machine that truly cannot reach out
I now do the final test with the machine on an isolated switch or with a firewall rule that blocks all outbound traffic, not just with Wi-Fi off. Then I go through every feature the users will actually touch: chat, uploading each type of document they will use, any search or retrieval over their files, restarting the machine, restarting the containers. I leave the logs open and look for anything mentioning a timeout, a DNS failure or a download. Every one of those is a dependency I missed.
On the engineering firm’s box, that test turned up one more surprise I had not hit in the lab: after a reboot, the front end took almost two minutes to start because it was waiting on a network check to time out. One environment variable fixed it.
Living with an offline model
Getting it running is half the job. Keeping it useful without a connection brings a few habits:
- Plan updates as deliveries. New model versions, runtime updates and security fixes all arrive on a drive now. The firm does this quarterly. I keep a written list of exact versions on the box so the next update is a known diff, not an archaeology project.
- The model’s knowledge is frozen. Every model has a training cutoff, and offline there is no web search to cover the gap. For the lab, that did not matter: the questions were about their own manuals, which were loaded as documents. For general questions about current events or new software versions, an offline model will be confidently out of date.
- Documentation has to live on the box too. When something breaks on an air-gapped machine, you cannot look up the error. I keep the runtime’s documentation and my own setup notes on the machine itself.
- Keep a fallback. The USB drive with llama.cpp and one known-good model goes with me on every visit. If the fancy front end breaks after an update, engineers still have a working assistant within five minutes.
Where offline makes sense
Besides secured networks like that lab, the offline question comes up more often than people expect: a research station or ship with an expensive satellite link, a field team working in areas with no coverage, a cabin with intermittent power and no broadband, or simply a household that wants its assistant to keep working when the internet is down. The pattern is the same in every case. The model is the easy part. The dependencies are the job.
The short version
A local LLM can run with no internet and no account, indefinitely. The model file and the runtime need nothing from the outside world. What needs the internet is the convenient tooling around them: registry pulls, helper models downloaded on first use, container images, packages and update checks. Stage all of it on a connected machine, verify the files, tell the software it is offline, and test on a machine that genuinely cannot reach out, using every feature your users will touch. My first offline box failed on its first PDF. The second one has not needed a network in eleven months.