Providers
Caudra talks to LLM providers over their HTTP APIs. Model jobs decide which configured or discovered model serves each kind of work.
Model jobs
Section titled “Model jobs”Caudra routes work through nine jobs: Chat, Plan, Subagent, Compact, Title, Goal, Extract, Fast, and Best. A global binding can pin a job to an exact provider/model-id or make it follow Chat, Plan, Fast, or Best. Explicit bindings report an error when their model is unavailable or disallowed.
/model opens a Jobs overview and the model list. Selecting a model on this page changes Chat. Selecting the Chat row jumps to its current model. Select another job to open its assignment page, press Esc to return to the overview, and use uppercase R to clear the open job's binding. /goal-model opens Goal directly. Jobs are opened from the overview rather than cycled with Tab.
Bindings are saved globally in the model.purposes row of Caudra's SQLite state database and apply across sessions. Unbound jobs use these rules:
| Job | Default when unbound |
|---|---|
| Chat | The anchor model |
| Plan | The anchor model |
| Subagent | The model currently running its parent |
| Compact | The model currently running the caller |
| Title | Fast |
| Goal | Fast |
| Extract | Fast |
| Fast | Provider fast config, curated preferred small model, cheapest priced model, fewest-parameter model, then the anchor |
| Best | Provider best config, curated flagship, then the anchor |
The anchor is the selected Chat model when a main turn starts. A Plan binding can select a distinct model, which Caudra uses for main turns sent in Plan mode. An explicit global Subagent binding overrides parent inheritance. A prompt profile's subagent_model overrides the global Subagent binding for tasks using that profile. See System Prompt Profiles.
See Sessions for Title, Completion goals for Goal, and Requirements for Extract.
Supply metadata
Section titled “Supply metadata”Provider catalogs record whether a model is small and whether it is the preferred model in its size lane. The tables below and caudra models render those facts with three markers:
| Marker | Meaning |
|---|---|
| Small | A small alternative |
| Fast | The preferred small model |
| Best | The provider's flagship |
A known non-small alternative and a model with no supply facts both have no marker. Markers describe provider supply. They are not capability tiers, and price or list order never creates one. Fast may still use price or parameter count as its final provider fallback without adding a marker.
Aggregators such as OpenRouter borrow supply facts from the upstream vendor entry in a vendor-prefixed model id. Tool deferral also uses the small fact: small and unknown models defer on-demand tools, while known non-small models receive them upfront. See Tools loaded on demand.
Provider purposes entries in providers.toml define the same supply facts. fast entries are small and best entries are non-small. The first entry in each list is preferred, so it receives the Fast or Best marker. Later fast entries receive Small, while later best entries remain unmarked. Entries match by prefix, so an endpoint serving a family of fine-tunes needs one line rather than one per variant:
[my-server.purposes]fast = "qwen3.8-27b" # covers qwen3.8-27b, qwen3.8-27b-canary, qwen3.8-27b-math7, ...The first entry also names the model that wins automatic resolution, so it has to name a model the endpoint actually serves. A prefix that matches nothing live would send requests to an id that does not exist.
Auth Reloading
Section titled “Auth Reloading”Caudra re-reads auth from storage and environment variables each time a new agent spawns (/new, retry, session load). If you run caudra auth login in another terminal or change an env var, the next session picks it up without a restart.
You can set multiple API keys in one env var (ANTHROPIC_API_KEY=sk-1,sk-2,sk-3) and they rotate automatically on rate-limit or auth errors.
Base URL Overrides
Section titled “Base URL Overrides”Most providers, custom ones included, honor a <SLUG>_BASE_URL env var, where <SLUG> is the slug in capitals with _ for - (anthropic -> ANTHROPIC_BASE_URL, openrouter -> OPENROUTER_BASE_URL). Set it to the origin of a proxy or a compatible endpoint and Caudra appends the API paths itself:
ANTHROPIC_BASE_URL=https://my-proxy.internal caudraIt wins over providers.toml and built-in defaults. ANTHROPIC_BASE_URL and OPENAI_BASE_URL are the same names the official SDKs use, so an existing proxy setup carries over as is. Three exceptions apply. ANTHROPIC_BASE_URL never receives Claude subscription tokens. OPENAI_BASE_URL only redirects the platform API, never the ChatGPT Coding Plan backend. XAI_BASE_URL only redirects the public API-key endpoint, never the OAuth CLI proxy.
You can also set base_url for a built-in provider in ~/.config/caudra/providers.toml. It overrides the built-in default and loses to the env var above:
[openai]base_url = "http://xxxx:1234/v1"The built-in provider still owns the slug, so protocol, api_key_env, discover_models and models are ignored with a warning. Use a custom slug if you need those.
Ollama reads OLLAMA_HOST and llama.cpp reads LLAMA_CPP_HOST. Neither reads <SLUG>_BASE_URL. A base_url in providers.toml wins over the host variable, and Caudra appends /v1 to either, so leave it off. Aperture reads APERTURE_HOST, which wins over the file. Copilot asks GitHub for the API endpoint of your account and ignores both settings.
Built-in Providers
Section titled “Built-in Providers”Anthropic
Section titled “Anthropic”- Env var:
ANTHROPIC_API_KEY(also supports subscription OAuth viacaudra auth login anthropic) - API:
https://api.anthropic.com/v1/messages - Features: Prompt caching, thinking mode (adaptive/budgeted), advanced tool use
| Marker | Models | Pricing (in/out per 1M tokens) | Context |
|---|---|---|---|
| Fast | claude-haiku-4-5 (default) | $1.00 / $5.00 | 200K ctx / 64K out |
| claude-sonnet-4-5 | $3.00 / $15.00 | 200K ctx / 64K out | |
| claude-sonnet-4-6 | $3.00 / $15.00 | 372K ctx / 64K out | |
| claude-sonnet-5-5 | $2.00 / $10.00 | 372K ctx / 128K out | |
| claude-sonnet-5 | $2.00 / $10.00 | 372K ctx / 128K out | |
| claude-sonnet-4 | $3.00 / $15.00 | 200K ctx / 64K out | |
| claude-opus-4-5 | $5.00 / $25.00 | 200K ctx / 64K out | |
| claude-opus-4-6 | $5.00 / $25.00 | 372K ctx / 128K out | |
| claude-opus-4-7 | $5.00 / $25.00 | 372K ctx / 128K out | |
| claude-opus-4-8 | $5.00 / $25.00 | 372K ctx / 128K out | |
| Best | claude-opus-5-5 (default) | $4.00 / $20.00 | 372K ctx / 128K out |
| claude-opus-5 | $5.00 / $25.00 | 372K ctx / 128K out | |
| claude-fable-5-1 | $10.00 / $50.00 | 372K ctx / 128K out | |
| claude-fable-5 | $10.00 / $50.00 | 372K ctx / 128K out | |
| claude-opus-4-0, claude-opus-4-1 | $15.00 / $75.00 | 200K ctx / 32K out |
Routing defaults: claude-haiku-4-5 (Fast), claude-opus-5-5 (Best)
Run caudra auth login anthropic to sign in to a Claude subscription through browser OAuth. Caudra stores the tokens in its state directory, refreshes them automatically, and shows subscription limits through /usage. Subscription requests always go to api.anthropic.com, even when ANTHROPIC_BASE_URL is set.
This experimental flow uses Claude Code's public client registration. Anthropic limits Pro and Max subscription tokens to official clients in its terms. The flow may stop working when Anthropic changes its OAuth or request protocol.
Recent Claude models accept up to 1M tokens. Caudra runs them at a 372k working window, which keeps cost and latency bounded. That window is an input budget: the model's output allowance sits on top of it rather than inside it, so Caudra holds back less of it before compaction. Add -1m to a model id, like claude-sonnet-4-6-1m, to open the full 1M window instead. Set context_window in providers.toml to pick any other size.
Amazon Bedrock
Section titled “Amazon Bedrock”If you already use Claude through AWS Bedrock, you can point Caudra at it instead of the direct Anthropic API. Set CLAUDE_CODE_USE_BEDROCK=1 and Caudra will route all Anthropic requests through Bedrock. The same models, the same features, just a different door.
You will need AWS_REGION and one of the following for auth:
| Method | Env vars |
|---|---|
| IAM credentials | AWS_ACCESS_KEY_ID + AWS_SECRET_ACCESS_KEY (and optionally AWS_SESSION_TOKEN) |
| Credentials file | AWS_PROFILE (defaults to default), reads ~/.aws/credentials |
| Bearer token | AWS_BEARER_TOKEN_BEDROCK |
| Gateway proxy | CLAUDE_CODE_SKIP_BEDROCK_AUTH=1 + ANTHROPIC_BEDROCK_BASE_URL (skips signing, useful behind a proxy that handles auth) |
You can override the model with ANTHROPIC_MODEL and the endpoint with ANTHROPIC_BEDROCK_BASE_URL. These env var names match Claude Code, so if you were already using Bedrock there, the same setup works here.
OpenAI
Section titled “OpenAI”- Env var:
OPENAI_API_KEY(also supports OAuth device flow) - API:
https://api.openai.com/v1
| Marker | Models | Pricing (in/out per 1M tokens) | Context |
|---|---|---|---|
| Fast | gpt-6-luna (gpt-6 default) | $0.10 / $0.50 | 372K ctx / 128K out |
| Fast | gpt-5.6-luna (gpt-5.6 default) | $0.20 / $1.20 | 372K ctx / 128K out |
| Small | gpt-5.4-nano | $0.20 / $1.25 | 400K ctx / 128K out |
| Small | gpt-5.4-mini | $0.75 / $4.50 | 400K ctx / 128K out |
| Small | gpt-4.1-nano | $0.10 / $0.40 | 1047K ctx / 32K out |
| Best | gpt-6-astra (gpt-6 default) | $10.00 / $50.00 | 372K ctx / 128K out |
| gpt-6.1-sol | $2.00 / $10.00 | 372K ctx / 128K out | |
| gpt-6-sol | $2.00 / $10.00 | 372K ctx / 128K out | |
| gpt-5.6-terra | $2.00 / $12.00 | 372K ctx / 128K out | |
| Best | gpt-5.6-sol (gpt-5.6 default) | $4.00 / $20.00 | 372K ctx / 128K out |
| gpt-4.1-mini | $0.40 / $1.60 | 1047K ctx / 32K out | |
| gpt-4.1 | $2.00 / $8.00 | 1047K ctx / 32K out | |
| o4-mini | $1.10 / $4.40 | 200K ctx / 100K out | |
| gpt-5.5 | $5.00 / $30.00 | 1050K ctx / 128K out | |
| gpt-5.4 | $2.50 / $15.00 | 1050K ctx / 128K out | |
| o3 | $2.00 / $8.00 | 200K ctx / 100K out | |
| gpt-5.3-codex | $1.75 / $14.00 | 400K ctx / 128K out | |
| gpt-5.2-codex | $1.75 / $14.00 | 400K ctx / 128K out | |
| gpt-5.2 | $1.75 / $14.00 | 400K ctx / 128K out | |
| gpt-5.1-codex-mini | $0.25 / $2.00 | 400K ctx / 128K out | |
| gpt-5.1-codex-max | $1.25 / $10.00 | 400K ctx / 128K out | |
| gpt-5.1-codex | $1.25 / $10.00 | 400K ctx / 128K out |
Routing defaults: gpt-6-luna (Fast), gpt-6-astra (Best)
A lane answers inside the release line you are on. On gpt-6 that is gpt-6-luna (Fast) and gpt-6-astra (Best). On gpt-5.6 that is gpt-5.6-luna (Fast) and gpt-5.6-sol (Best).
- Env var:
GEMINI_API_KEY - API:
https://generativelanguage.googleapis.com/v1beta - Features: Native Gemini API with thinking support
| Marker | Models | Pricing (in/out per 1M tokens) | Context |
|---|---|---|---|
| Fast | gemini-2.0-flash-lite (default) | $0.07 / $0.30 | 1048K ctx / 65K out |
| Best | gemini-2.5-pro (default) | $1.25 / $10.00 | 1048K ctx / 65K out |
| gemini-2.5-flash | $0.30 / $2.50 | 1048K ctx / 65K out |
Routing defaults: gemini-2.0-flash-lite (Fast), gemini-2.5-pro (Best)
Copilot
Section titled “Copilot”- Env var:
GH_COPILOT_TOKEN(or runcaudra auth login copilotto import a token from gh CLI, the Copilot client, or the system keyring) - API:
https://api.githubcopilot.com (or GraphQL-discovered Copilot API endpoint) - Features: Native Copilot Chat HTTP API with model endpoint discovery
| Marker | Models | Pricing (in/out per 1M tokens) | Context |
|---|---|---|---|
| Small | gpt-5-mini | $0.25 / $2.00 | 200K ctx / 100K out |
| Small | gpt-5.4-mini | $0.75 / $4.50 | 200K ctx / 100K out |
| Small | gpt-5.4-nano | $0.20 / $1.25 | 200K ctx / 100K out |
| Small | claude-haiku-4.5 | $1.00 / $5.00 | 200K ctx / 64K out |
| Small | gemini-3.5-flash | $1.50 / $9.00 | 200K ctx / 65K out |
| Small | mai-code-1-flash-picker | $0.75 / $4.50 | 200K ctx / 100K out |
| Small | gpt-6-luna | $0.10 / $0.50 | 200K ctx / 100K out |
| Fast | gpt-5.6-luna (default) | $0.20 / $1.20 | 200K ctx / 100K out |
| gemini-3.6-flash | $0.75 / $3.75 | 200K ctx / 65K out | |
| gemini-3.7-flash | $0.75 / $3.75 | 200K ctx / 65K out | |
| claude-sonnet-4.5, claude-sonnet-4.6 | $3.00 / $15.00 | 200K ctx / 64K out | |
| claude-sonnet-5.5 | $2.00 / $10.00 | 200K ctx / 100K out | |
| claude-sonnet-5 | $2.00 / $10.00 | 200K ctx / 100K out | |
| gpt-5.5 | $5.00 / $30.00 | 200K ctx / 100K out | |
| kimi-k2.7-code | $0.95 / $4.00 | 200K ctx / 100K out | |
| kimi-k3 | $3.00 / $15.00 | 200K ctx / 100K out | |
| gemini-3.1-pro-preview | $2.00 / $12.00 | 200K ctx / 65K out | |
| gpt-6.1-sol | $2.00 / $10.00 | 200K ctx / 100K out | |
| gpt-6-sol | $2.00 / $10.00 | 200K ctx / 100K out | |
| gpt-5.4 | $2.50 / $15.00 | 200K ctx / 100K out | |
| gpt-5.6-sol | $4.00 / $20.00 | 200K ctx / 100K out | |
| gpt-5.6-terra | $2.00 / $12.00 | 200K ctx / 100K out | |
| gpt-5.3-codex | $1.75 / $14.00 | 200K ctx / 100K out | |
| claude-opus-5.5 | $4.00 / $20.00 | 200K ctx / 128K out | |
| Best | claude-opus-5, claude-opus-4.8, claude-opus-4.7, claude-opus-4.6, claude-opus-4.5 (default) | $5.00 / $25.00 | 200K ctx / 64K out |
| claude-opus-4.8-fast, claude-fable-5 | $10.00 / $50.00 | 200K ctx / 100K out | |
| grok-4.5 | $2.00 / $6.00 | 200K ctx / 100K out | |
| grok-4.6 | $2.00 / $6.00 | 200K ctx / 100K out |
Routing defaults: gpt-5.6-luna (Fast), claude-opus-5 (Best)
Ollama
Section titled “Ollama”- Env var:
OLLAMA_HOSTfor local/remote (e.g.http://localhost:11434),OLLAMA_API_KEYfor auth - API:
http://localhost:11434/v1 - Features: Local or remote inference via OLLAMA_HOST, cloud fallback via OLLAMA_API_KEY
This provider talks the OpenAI-compatible /v1 API, so it also works with llama.cpp's server, LocalAI, or anything else that speaks the same protocol. Just point OLLAMA_HOST to the right address (e.g. http://localhost:8080 for llama.cpp).
LlamaCpp
Section titled “LlamaCpp”- Env var:
LLAMA_CPP_API_KEY - API:
http://localhost:8080/v1 - Features: Local or remote inference via LLAMA_CPP_HOST, set optional key via LLAMA_CPP_API_KEY
Connects to any OpenAI-compatible /v1 endpoint. Set LLAMA_CPP_HOST to your server address, such as http://localhost:8080, or run caudra auth login llama-cpp, which offers that address and saves your answer to providers.toml. Without either, Caudra reports that LLAMA_CPP_HOST is not set.
Mistral
Section titled “Mistral”- Env var:
MISTRAL_API_KEY - API:
https://api.mistral.ai/v1
| Marker | Models | Pricing (in/out per 1M tokens) | Context |
|---|---|---|---|
| Fast | ministral-14b-latest, ministral-14b-2512 (default) | $0.20 / $0.20 | 262K ctx |
| Best | mistral-medium-latest, mistral-medium-3.5, mistral-medium-3-5, mistral-medium-2604 (default) | $1.50 / $7.50 | 262K ctx |
| glm-5-2, zai-glm-5-2 | $1.40 / $4.40 | 1000K ctx | |
| mistral-small-latest, mistral-small-2603 | $0.15 / $0.60 | 262K ctx |
Routing defaults: ministral-14b-latest (Fast), mistral-medium-latest (Best)
- Env var:
ZHIPU_API_KEY(shared across both endpoints) - API endpoints:
https://api.z.ai/api/paas/v4https://api.z.ai/api/coding/paas/v4
| Marker | Models | Pricing (in/out per 1M tokens) | Context |
|---|---|---|---|
| Fast | glm-4.7-flash (default) | $0.00 / $0.00 | 200K ctx / 131K out |
| Small | glm-4.5-flash | $0.00 / $0.00 | 131K ctx / 98K out |
| Small | glm-4.5-air | $0.20 / $1.10 | 131K ctx / 98K out |
| Best | glm-5-code (default) | $1.20 / $5.00 | 200K ctx / 131K out |
| glm-5.2 | $1.00 / $3.20 | 1000K ctx / 131K out | |
| glm-5.1, glm-5 | $1.00 / $3.20 | 200K ctx / 131K out | |
| glm-4.7, glm-4.6 | $0.60 / $2.20 | 200K ctx / 131K out | |
| glm-4.5 | $0.60 / $2.20 | 131K ctx / 98K out |
Routing defaults: glm-4.7-flash (Fast), glm-5-code (Best)
DeepSeek
Section titled “DeepSeek”- Env var:
DEEPSEEK_API_KEY - API:
https://api.deepseek.com - Features: Thinking mode toggle (on/off), open-weight models
- Peak pricing: the prices below are off-peak; each turn is billed as it happens, at 2x during 01:00-04:00, 06:00-10:00 UTC
| Marker | Models | Pricing (in/out per 1M tokens) | Context |
|---|---|---|---|
| deepseek-v4-flash | $0.22 / $0.66 | 1000K ctx / 384K out | |
| Best | deepseek-v4-pro (default) | $0.66 / $1.98 | 1000K ctx / 384K out |
Routing defaults: deepseek-v4-pro (Best)
OpenRouter
Section titled “OpenRouter”- Env var:
OPENROUTER_API_KEY - API:
https://openrouter.ai/api/v1 - Features: 300+ models from all providers, prompt caching, provider routing
OpenRouter aggregates models from many providers behind a single API key. Browse available models at openrouter.ai/models. Use any model ID directly (e.g. openrouter/anthropic/claude-sonnet-4).
Synthetic
Section titled “Synthetic”- Env var:
SYNTHETIC_API_KEY - API:
https://api.synthetic.new/openai/v1 - Features: Reasoning effort support (low/medium/high), open-weight models
| Marker | Models | Pricing (in/out per 1M tokens) | Context |
|---|---|---|---|
| Fast | hf:zai-org/GLM-4.7-Flash (default) | $0.10 / $0.50 | 200K ctx / 131K out |
| Best | hf:moonshotai/Kimi-K2.5 (default) | $0.45 / $3.40 | 200K ctx / 131K out |
| hf:deepseek-ai/DeepSeek-V3.2 | $0.56 / $1.68 | 200K ctx / 131K out |
Routing defaults: hf:zai-org/GLM-4.7-Flash (Fast), hf:moonshotai/Kimi-K2.5 (Best)
TensorX
Section titled “TensorX”- Env var:
TENSORX_API_KEY - API:
https://api.tensorx.ai/v1 - Features: Open-weight models, zero data retention, prompt caching
No hardcoded model catalog. Use any model ID supported by this provider.
Opencode Zen
Section titled “Opencode Zen”- Env var:
OPENCODE_API_KEY - API:
https://opencode.ai/zen/v1 - Features: Dynamically discovered models via models.dev + all the models provided by Opencode Zen API
No hardcoded model catalog. Use any model ID supported by this provider.
By default Caudra hides free models from the Opencode catalog. To list free models (they use a public fallback, no API key needed), add this to ~/.config/caudra/providers.toml:
[opencode]enable_free_models = trueThe default is false.
- Env var:
XAI_API_KEY(also supports OAuth viacaudra auth login xai) - API endpoints:
https://api.x.ai/v1https://cli-chat-proxy.grok.com/v1
- Features: OAuth login, account-specific model catalog, Grok reasoning (low/medium/high/xhigh)
| Marker | Models | Pricing (in/out per 1M tokens) | Context |
|---|---|---|---|
| Best | grok-4.6 (default) | $2.00 / $6.00 | 500K ctx / 131K out |
| grok-4.5 | $2.00 / $6.00 | 500K ctx / 131K out | |
| grok-4.3 | $1.25 / $2.50 | 1000K ctx / 131K out |
Routing defaults: grok-4.6 (Best)
OAuth uses the same first-party xAI client as the official Grok CLI (caudra auth login xai). Browser login (PKCE) is the desktop default; device code is recommended over SSH or in a container. Tokens refresh automatically. After login, Caudra fetches your account catalog from GET /v1/models-v2 on the Grok CLI proxy and caches it for 15 minutes. XAI_BASE_URL only redirects the public API-key endpoint, never the OAuth proxy.
If ~/.grok/auth.json already exists, login offers to reuse it without writing that file.
Aperture
Section titled “Aperture”- Env var:
APERTURE_HOST(e.g.https://your-host.tailnet.ts.net) - API:
Aperture gateway (set APERTURE_HOST) - Features: Tailscale Aperture LLM gateway; set APERTURE_HOST or configure in providers.toml
Aperture discovers models from your gateway. Set APERTURE_HOST to your Tailscale Aperture endpoint (e.g. https://your-host.tailnet.ts.net). No API key needed, Tailscale handles auth.
Opencode Go
Section titled “Opencode Go”- Env var:
OPENCODE_API_KEY - API:
https://opencode.ai/zen/go/v1 - Features: Dynamically discovered models via models.dev + all the models provided by Opencode Go API
No hardcoded model catalog. Use any model ID supported by this provider. An API key is required.
Model Identifiers
Section titled “Model Identifiers”Models are referenced as provider/model_id:
anthropic/claude-sonnet-4-6openai/gpt-4.1xai/grok-4.6zai/glm-4.7If the model name is unique across providers, the prefix can be omitted.
providers.toml
Section titled “providers.toml”providers.toml lives in the config directory (~/.config/caudra/providers.toml on Linux/macOS, %APPDATA%\caudra\providers.toml on Windows). It is the file for provider overrides and custom HTTP providers. Two jobs:
- Tweak a built-in (pick a plan, change its base URL, set
enable_free_modelsfor Opencode). - Declare a custom provider that speaks OpenAI, Anthropic, or Google wire format.
# Point a built-in at a proxy. Env vars still win over this file.[anthropic]base_url = "https://my-proxy.internal"
# Full custom provider. Slug becomes the `provider/` prefix in model specs.[my-proxy]display_name = "My Proxy"protocol = "openai" # openai | openai-responses | anthropic | googlebase_url = "https://llm.example.com/v1"api_key_env = "MY_PROXY_API_KEY"default_model = "my-proxy/fast-v1"discover_models = true # also list models via the provider's /models endpoint
[my-proxy.purposes]fast = "fast-v1" # a prefix: also covers fast-v1-turbo, fast-v1-lora, ...best = ["smart-v1", "smart-v0"] # a list when one prefix cannot span them
[[my-proxy.models]]id = "fast-v1"context_window = 128000max_output_tokens = 16384pricing_input = 0.5pricing_output = 1.5
[[my-proxy.models]]id = "smart-v1"context_window = 200000max_output_tokens = 32000supports_thinking = truesupports_vision = falseThe file can start with version = 1, and a file without it counts as version 1. Caudra writes the key whenever it saves the file. A newer version stops Caudra with an error rather than being misread. Because version belongs to the file, a custom provider cannot use it as a name. See Config file versions.
caudra config example providers prints every providers.toml key with its default, all commented out. Reference configs shows the same text.
Provider fields
Section titled “Provider fields”| Field | Type | Default | Env | Description |
|---|---|---|---|---|
display_name | string | the built-in name, or the slug | - | The name pickers and auth status show |
protocol | string | required | - | The wire format: openai, openai-responses, anthropic, or google |
base_url | string | the plan URL, or the built-in URL | <SLUG>_BASE_URL | The API origin. Caudra appends the protocol paths |
plan | string | unset | - | A built-in plan key, which sets the base URL and the default model |
api_key_env | string | <SLUG>_API_KEY | - | The environment variable that holds the API key |
api_key | string | unset | - | An API key, stored as plain text. Caudra tries the environment variable and saved credentials first |
default_model | string | unset | - | The model to use after login when none is saved yet, such as my-provider/my-model |
discover_models | bool | false | - | Also list the models the provider's model endpoint reports |
enable_free_models | bool | unset | - | Opencode only. Show the free models of its catalog. Unset counts as false |
overrides | table | unset | - | Aperture only. Overrides for the upstream providers it routes, keyed by upstream id |
model_defaults | table | unset | - | Model keys for every model of the provider |
purposes | table | unset | - | Model id prefixes for the fast and best purposes |
models | table[] | unset | - | The models the provider serves |
A built-in slug keeps its compiled protocol, model catalog, and auth setup, so it ignores protocol, api_key_env, discover_models, models and enable_free_models. Opencode still reads enable_free_models.
A [SLUG.purposes] table takes these keys:
| Field | Type | Default | Description |
|---|---|---|---|
fast | string | string[] | unset | Model id prefixes for small, fast models, best first. A prefix covers every id that starts with it, and the first one also names the model that fills the slot, so it has to be a real id |
best | string | string[] | unset | Model id prefixes for flagship models, best first. A prefix cannot also be in fast. A model a job is bound to in the picker wins over both lists |
Model fields
Section titled “Model fields”Each [[SLUG.models]] entry declares one model:
| Field | Type | Default | Description |
|---|---|---|---|
id | string | required | The model id, which makes the spec SLUG/ID |
context_window | integer | discovered, or the protocol default | Tokens of context |
max_output_tokens | integer | discovered, or the protocol default | The most tokens one response may hold |
supports_tool_examples | bool | false | Send tool examples as a structured field. It is off unless declared, because the protocol says nothing about the model behind it |
supports_thinking | bool | discovered, or the protocol default | The model accepts extended thinking |
requires_thinking | bool | false | For an API that rejects requests with thinking off. It implies supports_thinking and raises thinking to minimal effort when it is off, compaction included |
supports_vision | bool | false | The model accepts images. When false, image input and view_image are off |
supports_pdf | bool | false | anthropic and openai-responses only. The model reads a PDF that webfetch attaches inside its tool result. When it is off, webfetch returns the text of the PDF instead |
supports_cache_breakpoints | bool | false | openai-responses only. The endpoint honours an explicit prompt_cache_breakpoint, so the system prompt closes with one |
reasoning_options | table[] | unset | The reasoning controls the model takes, such as [{ type = "effort", values = ["low", "high"] }]. A type is toggle, effort with values, or budget_tokens with an optional min and max. [] declares that it takes none, so Caudra sends no reasoning level |
pricing_input | float | 0 | USD per million input tokens |
pricing_output | float | 0 | USD per million output tokens |
pricing_cache_write | float | 0 | USD per million tokens written to the prompt cache |
pricing_cache_read | float | 0 | USD per million tokens read from the prompt cache |
pricing_fast_input | float | unset | USD per million input tokens in fast mode |
pricing_fast_output | float | unset | USD per million output tokens in fast mode |
Model defaults
Section titled “Model defaults”A models entry only applies to the exact id it names. When a provider's ids change often, or discover_models finds models you never declared, put the shared settings in model_defaults instead:
[my-proxy.model_defaults]context_window = 229376max_output_tokens = 32768reasoning_options = []
[[my-proxy.models]]id = "smart-v1"max_output_tokens = 64000It takes the same fields as a models entry apart from id, and applies to every model of the provider including discovered ones. A matching models entry wins field by field, so smart-v1 above keeps the 229376-token window and raises only its output cap. Anything a model neither declares nor inherits falls back to discovery, then to the protocol default.
reasoning_options = [] is a declaration, not an omission: it says the endpoint takes no reasoning controls, so Caudra sends no reasoning_effort. Leaving it unset instead lets a thinking level chosen for another model reach an endpoint that rejects it.
A models entry whose id matches no live model is logged once: every setting on it is ignored, which otherwise looks like Caudra disregarding the config.
Custom slugs must not reuse a built-in provider name. A bad TOML parse exits with code 2 at startup so a typo cannot silently empty the registry.
You can also create a custom provider interactively with caudra auth login and choosing the custom option. That writes a starter entry to this file.
Aperture overrides
Section titled “Aperture overrides”Aperture proxies upstream providers, exposing each model as aperture/<upstream>/<model>. Overrides keyed by upstream provider id live under [aperture.overrides]:
[aperture.overrides.llmserver]base = "llama-cpp"context_window = 131072max_output_tokens = 16384
[aperture.overrides.llmserver.models."qwen-3.6"]context_window = 262144supports_vision = trueProvider-level fields apply to every model from that upstream. Per-model entries under models win field by field, and take the same keys apart from models. Model ids containing dots must be quoted ("qwen3.6") since TOML treats a bare dotted key as a nested table.
| Field | Type | Default | Description |
|---|---|---|---|
context_window | integer | unset | Tokens of context |
max_output_tokens | integer | unset | The most tokens one response may hold |
supports_thinking | bool | unset | The models accept extended thinking |
supports_vision | bool | unset | The models accept images |
base | string | unset | The native provider an opaque upstream works like, such as llama-cpp, google, or anthropic. Caudra warns about a value it does not know and ignores it |
path_prefix | string | /v1, /v1beta for Gemini routes, none for Anthropic and Z.AI | The path Caudra sends ahead of each request, which Aperture appends to the upstream base URL. Set it to "" when that URL already has its own path |
models | table | {} | Overrides for single models, keyed by model id, which win key by key. Quote an id that holds a dot, such as models."qwen-3.6" |
Caudra sends /v1 (or /v1beta for Gemini routes, nothing for Anthropic and Z.AI), and Aperture appends that path to the upstream's base url. If an upstream base url already carries its own path, set path_prefix = "" for it to avoid a doubled path. Z.AI defaults to no prefix since its API path has no /v1 segment; point the upstream base url at the full API root (e.g. https://api.z.ai/api/paas/v4).
Some built-ins ship multiple plans (different base URLs or default models). caudra auth login <provider> asks which plan to use when more than one exists. You can also set it in TOML:
[mistral]plan = "coding"
[zai]plan = "coding"Current plans:
| Provider | Plan | What it does |
|---|---|---|
| Mistral | standard | Standard at https://api.mistral.ai/v1, default mistral/mistral-medium-latest |
| Mistral | coding | Vibe / Coding at https://api.mistral.ai/v1, default mistral/mistral-vibe-cli-latest |
| Z.AI | standard | Pay-as-you-go at https://api.z.ai/api/paas/v4, default zai/glm-5.1 |
| Z.AI | coding | Coding plan at https://api.z.ai/api/coding/paas/v4, default zai/glm-5-code |
Env <SLUG>_BASE_URL still wins over both the plan and a base_url in this file.
Dynamic Providers
Section titled “Dynamic Providers”To add a custom provider or proxy, drop an executable script into the config providers/ directory (~/.config/caudra/providers/ on Linux/macOS, %APPDATA%\caudra\providers\ on Windows). The script must handle these subcommands:
| Subcommand | Timeout | What it does |
|---|---|---|
info | 5s | Return JSON with display_name, base provider, has_auth |
models | 5s | Return JSON array of model entries (optional) |
resolve | 30s | Return auth JSON (base_url, headers) |
login | interactive | OAuth or credential flow |
logout | interactive | Clear credentials |
refresh | 30s | Refresh auth tokens |
resolve is called each time a new agent spawns, so scripts should read tokens from disk instead of caching them in memory. That way auth changes from other processes get picked up.
The base field specifies which built-in provider to inherit the model catalog from. Valid values: anthropic, openai, google, copilot, ollama, llama-cpp, mistral, zai, deepseek, openrouter, synthetic, tensorx, opencode, xai, aperture.
If your provider serves models not in the base catalog, add a models subcommand returning:
[{"id": "my-model-v2", "context_window": 200000, "max_output_tokens": 16384}]Only id is required. Optional fields: context_window (128K), max_output_tokens (16K), pricing ({input, output, cache_write, cache_read}, all per 1M tokens), supports_tool_examples (defaults to the base provider's setting), supports_thinking (defaults to the base provider's setting), requires_thinking (default false; for APIs that reject requests with thinking off, raises it to minimal effort and implies supports_thinking), supports_vision (defaults to the base provider's setting; when false, image input and the view_image tool are disabled). Without this subcommand, the base provider's models are used.
A llama-cpp model can replace Caudra's token-budget mapping with its native thinking fields. Each thinking mode maps to a JSON fragment merged into the request body:
[{ "id": "reasoning-model", "supports_thinking": true, "thinking_fields": { "off": {"reasoning_effort": "none"}, "adaptive": {"reasoning_effort": "medium"}, "low": {"reasoning_effort": "low"}, "medium": {"reasoning_effort": "medium"}, "xhigh": {"reasoning_effort": "xhigh"} }}]off is used when thinking is off, adaptive when thinking is on without a chosen level. Any other key is an effort level, one of none, minimal, low, medium, high, xhigh, max. The levels you declare are the ones the model accepts: whatever you ask for snaps into them, downwards first, so a level the model never advertised is never sent. Every part is optional.
Fragments are merged into the body, so nesting works too. A template toggle is just a fragment:
"thinking_fields": { "off": {"chat_template_kwargs": {"enable_thinking": false}}, "adaptive": {"chat_template_kwargs": {"enable_thinking": true}}}Named modes send only these fields, no token budget. An explicit /thinking <budget> snaps into the levels you declared; a model that declares none gets the adaptive fragment plus thinking_budget_tokens. Any mode you left undeclared falls back to the usual thinking_budget_tokens mapping, so no request ever ends up saying nothing. Models without thinking_fields keep the existing llama.cpp behavior.
Dynamic provider models are namespaced as {slug}/{model_id} (e.g. myproxy/claude-sonnet-4-6).
Script Name Rules
Section titled “Script Name Rules”- Must start with a letter or digit
- Only letters, digits, underscores, and hyphens after that
- Can't reuse a built-in provider's slug
- Must be executable