Run local AI model inference with Ollama and LM Studio
Run model requests on your own hardware with no API bill and no internet required for inference. With sync, online tools, master telemetry and AI observability (including replay) off, prompts and responses stay on your machine. Here is how to set up Ollama or LM Studio.
The More AI team · Updated September 14, 2026
Why run a model locally
A local endpoint removes the hosted-model hop only when the chosen model actually runs on your machine. A localhost address can still route a request to a cloud model, so check the selected model and runtime settings.
Local inference has no provider per-token bill. Hardware, electricity, storage and your time still have a cost. Capability depends on the exact model, quantization, context length and available memory; test one small task before choosing it for a workflow.
Ollama: the quickest route
Start a local server before connecting More AI. Ollama uses http://localhost:11434/v1 by default. In LM Studio, load the model and start the local server, normally http://localhost:1234/v1; enter its token if you enabled server authentication. These local endpoints differ from the providers’ hosted services.
LM Studio: a friendly model manager
Ollama defaults to http://localhost:11434/v1; LM Studio to http://localhost:1234/v1. Start the server and load a model before connecting. Local Ollama is keyless; LM Studio can require a token when its optional server authentication is enabled. Custom-server protocol and tool support must be checked.
Any other server: custom endpoints
Running vLLM, llama.cpp or a LiteLLM proxy? Add a custom provider and point it at any endpoint that speaks the OpenAI-compatible Chat Completions API (the base URL usually ends in /v1) — or at an Anthropic-compatible Messages endpoint. That covers practically every self-hosted model server, on your machine or on a box in your LAN.

Mix local and cloud
You don’t have to choose once. Keys and local providers live side by side, and every conversation picks its own model — so a practical setup is a local model for drafts and quick questions, and a frontier model for harder tasks. Switching is two clicks, not a reconfiguration.
Cowork and Code work with local models through the same engine, and earlier Chat conversations can keep using them too. Prefer the most capable local model your hardware can hold: planning and tool use reward a stronger model.
Check the server before testing an agent
Ollama and LM Studio expose OpenAI-compatible interfaces, but compatibility does not imply that every model supports reliable tools. Use a model whose current model card and runtime support tool calling, then verify it with a small file task. A successful chat response is only the first check.
- Install the runtime and download a model that fits your hardware. Record runtime version, model ID, quantization, context setting and available RAM/VRAM.
- Start the local server: Ollama normally uses http://localhost:11434/v1; LM Studio normally uses http://localhost:1234/v1. Keep it bound to your machine unless you deliberately need network access.
- Enable the matching provider in Settings → Providers. Use the server token if LM Studio authentication is enabled. Refresh the model list and select the exact downloaded model.
- Send a short text request, then run the file task below. Record unsupported-tool errors and stop loops rather than treating repeated attempts as progress.
A small local file test
The kit contains fictional project notes, English/Russian/Chinese prompts and an expected answer. Ask the agent to read the notes and create output/summary.txt with three project names and the pending task T03. Do not allow input edits or web access for this exercise.
Open the output yourself. All three projects must be present and T03 must remain pending; the original notes must be unchanged. Record runtime/model versions, elapsed time, tool errors and manual corrections. This is a test recipe, not a published More AI benchmark.
Check data flow separately from model location
A localhost model address proves where the inference server is addressed, not where every app feature sends data. Signed-out desktop work is not synchronized. After sign-in, supported sync categories are on by default unless you turn them off; cloud execution is a separate switch. Cowork file E2E is optional and initially off.
For a local-only exercise, prepare the model and dependencies first, disable sync and cloud execution, online tools and remote connectors, master telemetry and AI observability including replay. Then disconnect the network and repeat the small task. A passing disconnected run confirms only that task under those settings; it is not a complete privacy audit or a promise that all features work offline.
Frequently asked questions
Is running a local model really free?
Local inference has no provider per-token bill. Hardware, electricity, storage and your time still have a cost. Capability depends on the exact model, quantization, context length and available memory; test one small task before choosing it for a workflow.
Does everything work offline?
Only a self-contained task whose model, dependencies and input files are already available locally can be tested offline. Sync, cloud execution, web search, remote connectors and other online services require a connection. Check telemetry and AI observability/replay separately; a disconnected test is not a complete privacy audit.
What hardware do I need?
Requirements depend on model size, quantization, context and concurrent work. Check the runtime and model requirements, leave room for the operating system, then record memory use and latency on a small task. No single RAM number guarantees useful agent performance.
Which local servers does More AI support?
Ollama defaults to http://localhost:11434/v1; LM Studio to http://localhost:1234/v1. Start the server and load a model before connecting. Local Ollama is keyless; LM Studio can require a token when its optional server authentication is enabled. Custom-server protocol and tool support must be checked.
Can the coding agent use a local model?
Yes — Code mode runs through the same engine, so a local model can plan and edit too. Agent work rewards stronger models, so use the most capable local model your hardware can hold, or switch that one session to a cloud key.