API Route
API Route is a hosted, OpenAI-compatible multi-model API. PrivateGPT can use its Chat Completions endpoint at https://global.api-route.com/v1 through the existing OpenAI-compatible integration.
This setup sends chat requests to API Route and keeps embeddings on a separate local Ollama server. It does not require a new PrivateGPT backend.
Setup
Create a key and choose a chat model
- Create an account, add balance, and create a key in API Keys.
- Set
OPENAI_API_KEYin your environment or.envfile. Keep the secret out of committed files. - List the models available to that key:
Set API_ROUTE_MODEL to the exact ID of a chat model from the response:
Replace the placeholder before running PrivateGPT. Availability depends on the key’s group and model permissions; the pricing page shows the current catalog. See the API access guide for the endpoint and authentication contract.
Start a local embeddings server
Start Ollama and pull the embeddings model:
These examples use Ollama’s OpenAI-compatible API at http://localhost:11434/v1, with the separate placeholder key ollama.
Run PrivateGPT
Package install
Docker
uv (local)
Follow the corresponding installation guide first, including PrivateGPT’s other required services.
Model limits and capabilities
- Use the model ID returned by
/v1/modelsunchanged; do not add an upstream provider prefix unless it is already part of that ID. - The catalog can include non-chat models. Choose a Chat Completions model for
PGPT_LLM_DEFAULT. - Streaming, tools, reasoning, and image input depend on the selected model and route. A catalog entry alone does not certify every capability.
- If discovery does not provide a context limit or capability metadata recognized by PrivateGPT, configure these explicitly in a model profile. Do not assume PrivateGPT’s fallback context window matches the route’s limit.
- The examples use local embeddings independently of the gateway’s embedding support. Retrieved document excerpts included in a chat prompt are still sent to the hosted API.

