> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.privategpt.dev/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.privategpt.dev/_mcp/server.

# Cheaper Inference

> Use Cheaper Inference as an OpenAI-compatible cloud gateway for models from OpenAI, Anthropic, Google, DeepSeek, Z.ai and other labs.

[Cheaper Inference](https://cheaperinference.com) is an OpenAI-compatible gateway at `https://api.cheaperinference.com/v1`. One API key gives access to models from OpenAI, Anthropic, Google, DeepSeek, Z.ai and other labs. Each model costs 15–60% less than the list price of its lab.

## Capabilities with PrivateGPT

| Capability                       | Status            |
| -------------------------------- | ----------------- |
| Model discovery (`/v1/models`)   | ✅                 |
| Tokenizer endpoint (`/tokenize`) | ❌                 |
| Embeddings                       | ❌                 |
| Tool / function calling          | ✅ model-dependent |
| Streaming                        | ✅                 |
| Vision / image input             | ✅ model-dependent |

The gateway has no embeddings endpoint. Use a second server for embeddings, for example [Ollama](/providers/ollama).

---

## Setup

#### Get an API key

1. Sign up at [cheaperinference.com](https://cheaperinference.com).
2. Create an API key in the dashboard.

#### Start an embeddings server

```bash
ollama pull mxbai-embed-large
```

The embeddings API is now available at `http://localhost:11434/v1`.

#### Run PrivateGPT

#### Package install

```bash
OPENAI_API_BASE=https://api.cheaperinference.com/v1 \
  OPENAI_API_KEY=your-cheaperinference-api-key \
  OPENAI_EMBEDDING_API_BASE=http://localhost:11434/v1 \
  PGPT_LLM_DEFAULT=gpt-5.4 \
  private-gpt serve
```

#### Docker

```bash
docker run -p 8080:8080 \
  -e OPENAI_API_BASE=https://api.cheaperinference.com/v1 \
  -e OPENAI_API_KEY=your-cheaperinference-api-key \
  -e OPENAI_EMBEDDING_API_BASE=http://host.docker.internal:11434/v1 \
  -e PGPT_LLM_DEFAULT=gpt-5.4 \
  zylonai/private-gpt:latest
```

#### uv (local)

```bash
OPENAI_API_BASE=https://api.cheaperinference.com/v1 \
  OPENAI_API_KEY=your-cheaperinference-api-key \
  OPENAI_EMBEDDING_API_BASE=http://localhost:11434/v1 \
  PGPT_LLM_DEFAULT=gpt-5.4 \
  uv run private-gpt serve
```

---

## Notes

* `/v1/models` returns `context_length` for each chat model, so PrivateGPT sets the context window automatically.
* `/v1/models` also lists image and video models. Set `PGPT_LLM_DEFAULT` to a chat model, such as `gpt-5.4`.