Skip to navigation

Cheaper Inference

Cheaper Inference is an OpenAI-compatible gateway at https://api.cheaperinference.com/v1. One API key gives access to models from OpenAI, Anthropic, Google, DeepSeek, Z.ai and other labs. Each model costs 15–60% less than the list price of its lab.

Capabilities with PrivateGPT

CapabilityStatus
Model discovery (/v1/models)✅
Tokenizer endpoint (/tokenize)❌
Embeddings❌
Tool / function calling✅ model-dependent
Streaming✅
Vision / image input✅ model-dependent

The gateway has no embeddings endpoint. Use a second server for embeddings, for example Ollama.


Setup

1

Get an API key

  1. Sign up at cheaperinference.com.
  2. Create an API key in the dashboard.
2

Start an embeddings server

ollama pull mxbai-embed-large

The embeddings API is now available at http://localhost:11434/v1.

3

Run PrivateGPT

OPENAI_API_BASE=https://api.cheaperinference.com/v1 \
OPENAI_API_KEY=your-cheaperinference-api-key \
OPENAI_EMBEDDING_API_BASE=http://localhost:11434/v1 \
PGPT_LLM_DEFAULT=gpt-5.4 \
private-gpt serve

Notes

  • /v1/models returns context_length for each chat model, so PrivateGPT sets the context window automatically.
  • /v1/models also lists image and video models. Set PGPT_LLM_DEFAULT to a chat model, such as gpt-5.4.