Skip to navigation

API Route

API Route is a hosted, OpenAI-compatible multi-model API. PrivateGPT can use its Chat Completions endpoint at https://global.api-route.com/v1 through the existing OpenAI-compatible integration.

This setup sends chat requests to API Route and keeps embeddings on a separate local Ollama server. It does not require a new PrivateGPT backend.

Setup

1

Create a key and choose a chat model

  1. Create an account, add balance, and create a key in API Keys.
  2. Set OPENAI_API_KEY in your environment or .env file. Keep the secret out of committed files.
  3. List the models available to that key:
curl https://global.api-route.com/v1/models \
-H "Authorization: Bearer $OPENAI_API_KEY"

Set API_ROUTE_MODEL to the exact ID of a chat model from the response:

export API_ROUTE_MODEL="your-chat-model-id"

Replace the placeholder before running PrivateGPT. Availability depends on the key’s group and model permissions; the pricing page shows the current catalog. See the API access guide for the endpoint and authentication contract.

2

Start a local embeddings server

Start Ollama and pull the embeddings model:

ollama pull mxbai-embed-large

These examples use Ollama’s OpenAI-compatible API at http://localhost:11434/v1, with the separate placeholder key ollama.

3

Run PrivateGPT

OPENAI_API_BASE=https://global.api-route.com/v1 \
OPENAI_EMBEDDING_API_BASE=http://localhost:11434/v1 \
OPENAI_EMBEDDING_API_KEY=ollama \
PGPT_LLM_DEFAULT="$API_ROUTE_MODEL" \
PGPT_EMBEDDING_DEFAULT=mxbai-embed-large \
private-gpt serve

Follow the corresponding installation guide first, including PrivateGPT’s other required services.

Model limits and capabilities

  • Use the model ID returned by /v1/models unchanged; do not add an upstream provider prefix unless it is already part of that ID.
  • The catalog can include non-chat models. Choose a Chat Completions model for PGPT_LLM_DEFAULT.
  • Streaming, tools, reasoning, and image input depend on the selected model and route. A catalog entry alone does not certify every capability.
  • If discovery does not provide a context limit or capability metadata recognized by PrivateGPT, configure these explicitly in a model profile. Do not assume PrivateGPT’s fallback context window matches the route’s limit.
  • The examples use local embeddings independently of the gateway’s embedding support. Retrieved document excerpts included in a chat prompt are still sent to the hosted API.