Use your own Ollama model with your own web search

Local mode is meant to be as private — and as friendly to use — as we can make it: prompts go to a model you run (typically on the Ollama engine), and conversations are saved only as Markdown files on this device (the TeeChat folder you choose). TeeChat does not host a chat archive on that path.

To make the AI more useful, web search is sometimes essential. Local mode cannot use Server execution; configure Client search instead (see section 4 below). The app will call your search URL and API key — Brave Search API, or any JSON search API you subscribe to.

The web app at chat.teechat.ai stays on hosted confidential inference. Use the desktop app for Local mode.

What goes where

PieceLocal inferenceConfidential (hosted)
Chat modelOllama / vLLM / LM Studio on this device or LANTeeChat confidential engines
Background helper (task model)You pick under Settings → Models & API → Background → On this deviceTeeChat hosted default (or your Hosted chat slot)
Web searchYou set Client URL + API key under Settings → Web search → Endpoints (runs on this device)TeeChat gateway search
Conversation filesYour TeeChat folder either waySame folder

Local inference is not OPE-encrypted. Prompts go only to the local service. Prompts are rewritten by the task model before they are sent to the search engine you configured — so searches are more targeted, and so less of your raw message is exposed through the search engine.

1. Install Ollama and pull a model

  1. Install Ollama from ollama.com and start the app.
  2. Pull a chat model that fits your RAM, for example:
ollama pull qwen2.5
ollama list
  1. Confirm the API answers:
curl -s http://127.0.0.1:11434/api/tags

You should see a JSON list of models. TeeChat also auto-detects vLLM (:8000), LM Studio (:1234), and llama.cpp if those are running instead.

2. Point TeeChat at the local engine

  1. Open the desktop app → Settings → Inference → On this device.
  2. Leave the server URL empty to auto-detect http://127.0.0.1:11434, or type that URL.
  3. Tap Test connection. You should see Connected — Ollama at http://127.0.0.1:11434.
  4. Choose On this device. The top bar also shows Local | Confidential once an engine is reachable.

Sign-in is not required for Local chat. New installs still default to Confidential; switch in Settings | Inference | On this device.

3. Pick a Local background helper (task model)

Search prep, chat titles, and folder suggestions use a background helper, not the chat model in the top bar. Local and Confidential each keep their own slot on this device.

  1. Optionally pull a smaller/faster tag for helper jobs (you can also reuse the chat model):
ollama pull qwen2.5:7b
  1. Open Settings → Models & API → Background.
  2. Expand Advanced: change background helpers.
  3. Under On this device, choose a model from your Ollama / vLLM list — not the hosted default. Leave Hosted chat alone unless you also want to change Confidential helpers.

If On this device is Not set — choose a model on this device…, Local web search cannot prepare queries (No task model is configured for search prep). TeeChat then asks before sending a redacted fallback — or skips search — depending on When task-model prep fails under Settings → Web search → Overview.

  1. Settings → Web search → Overview, set Execution mode to Client (browser calls search URL). Server (POST to configured URL) is not available in Local mode.
  2. Open Endpoints.
  3. Set Client URL and Search API key (merged as {{apiKey}}; stored on this device only) (POST JSON or GET template).
  4. Save search settings.
  5. In the composer, leave Web search on for threads that need sources.

If the Client URL is empty while Local is on, TeeChat skips search and tells you to configure an endpoint. It will not call api.teechat.ai / regional gateway search on your behalf.

The named template we ship is Brave Search API. Custom accepts any JSON API that returns { "results": [{ "title", "url", "snippet" }] } (or another shape TeeChat already parses).

5. Quick checklist

Privacy

Outbound search strings go through best-effort PII redaction. That is not a compliance guarantee.

For verified hosted engines (no local model), switch to Confidential and follow How to verify confidential chat.

Revision log

  1. Document Local background helper (task model) setup for search prep.

← All posts