Skip to content

Providers

View Markdown~8,892 tokens

Caudra talks to LLM providers over their HTTP APIs. Model jobs decide which configured or discovered model serves each kind of work.

Caudra routes work through nine jobs: Chat, Plan, Subagent, Compact, Title, Goal, Extract, Fast, and Best. A global binding can pin a job to an exact provider/model-id or make it follow Chat, Plan, Fast, or Best. Explicit bindings report an error when their model is unavailable or disallowed.

/model opens a Jobs overview and the model list. Selecting a model on this page changes Chat. Selecting the Chat row jumps to its current model. Select another job to open its assignment page, press Esc to return to the overview, and use uppercase R to clear the open job's binding. /goal-model opens Goal directly. Jobs are opened from the overview rather than cycled with Tab.

Bindings are saved globally in the model.purposes row of Caudra's SQLite state database and apply across sessions. Unbound jobs use these rules:

JobDefault when unbound
ChatThe anchor model
PlanThe anchor model
SubagentThe model currently running its parent
CompactThe model currently running the caller
TitleFast
GoalFast
ExtractFast
FastProvider fast config, curated preferred small model, cheapest priced model, fewest-parameter model, then the anchor
BestProvider best config, curated flagship, then the anchor

The anchor is the selected Chat model when a main turn starts. A Plan binding can select a distinct model, which Caudra uses for main turns sent in Plan mode. An explicit global Subagent binding overrides parent inheritance. A prompt profile's subagent_model overrides the global Subagent binding for tasks using that profile. See System Prompt Profiles.

See Sessions for Title, Completion goals for Goal, and Requirements for Extract.

Provider catalogs record whether a model is small and whether it is the preferred model in its size lane. The tables below and caudra models render those facts with three markers:

MarkerMeaning
SmallA small alternative
FastThe preferred small model
BestThe provider's flagship

A known non-small alternative and a model with no supply facts both have no marker. Markers describe provider supply. They are not capability tiers, and price or list order never creates one. Fast may still use price or parameter count as its final provider fallback without adding a marker.

Aggregators such as OpenRouter borrow supply facts from the upstream vendor entry in a vendor-prefixed model id. Tool deferral also uses the small fact: small and unknown models defer on-demand tools, while known non-small models receive them upfront. See Tools loaded on demand.

Provider purposes entries in providers.toml define the same supply facts. fast entries are small and best entries are non-small. The first entry in each list is preferred, so it receives the Fast or Best marker. Later fast entries receive Small, while later best entries remain unmarked. Entries match by prefix, so an endpoint serving a family of fine-tunes needs one line rather than one per variant:

[my-server.purposes]
fast = "qwen3.8-27b" # covers qwen3.8-27b, qwen3.8-27b-canary, qwen3.8-27b-math7, ...

The first entry also names the model that wins automatic resolution, so it has to name a model the endpoint actually serves. A prefix that matches nothing live would send requests to an id that does not exist.

Caudra re-reads auth from storage and environment variables each time a new agent spawns (/new, retry, session load). If you run caudra auth login in another terminal or change an env var, the next session picks it up without a restart.

You can set multiple API keys in one env var (ANTHROPIC_API_KEY=sk-1,sk-2,sk-3) and they rotate automatically on rate-limit or auth errors.

Most providers, custom ones included, honor a <SLUG>_BASE_URL env var, where <SLUG> is the slug in capitals with _ for - (anthropic -> ANTHROPIC_BASE_URL, openrouter -> OPENROUTER_BASE_URL). Set it to the origin of a proxy or a compatible endpoint and Caudra appends the API paths itself:

Terminal window
ANTHROPIC_BASE_URL=https://my-proxy.internal caudra

It wins over providers.toml and built-in defaults. ANTHROPIC_BASE_URL and OPENAI_BASE_URL are the same names the official SDKs use, so an existing proxy setup carries over as is. Three exceptions apply. ANTHROPIC_BASE_URL never receives Claude subscription tokens. OPENAI_BASE_URL only redirects the platform API, never the ChatGPT Coding Plan backend. XAI_BASE_URL only redirects the public API-key endpoint, never the OAuth CLI proxy.

You can also set base_url for a built-in provider in ~/.config/caudra/providers.toml. It overrides the built-in default and loses to the env var above:

[openai]
base_url = "http://xxxx:1234/v1"

The built-in provider still owns the slug, so protocol, api_key_env, discover_models and models are ignored with a warning. Use a custom slug if you need those.

Ollama reads OLLAMA_HOST and llama.cpp reads LLAMA_CPP_HOST. Neither reads <SLUG>_BASE_URL. A base_url in providers.toml wins over the host variable, and Caudra appends /v1 to either, so leave it off. Aperture reads APERTURE_HOST, which wins over the file. Copilot asks GitHub for the API endpoint of your account and ignores both settings.

  • Env var: ANTHROPIC_API_KEY (also supports subscription OAuth via caudra auth login anthropic)
  • API: https://api.anthropic.com/v1/messages
  • Features: Prompt caching, thinking mode (adaptive/budgeted), advanced tool use
MarkerModelsPricing (in/out per 1M tokens)Context
Fastclaude-haiku-4-5 (default)$1.00 / $5.00200K ctx / 64K out
claude-sonnet-4-5$3.00 / $15.00200K ctx / 64K out
claude-sonnet-4-6$3.00 / $15.00372K ctx / 64K out
claude-sonnet-5-5$2.00 / $10.00372K ctx / 128K out
claude-sonnet-5$2.00 / $10.00372K ctx / 128K out
claude-sonnet-4$3.00 / $15.00200K ctx / 64K out
claude-opus-4-5$5.00 / $25.00200K ctx / 64K out
claude-opus-4-6$5.00 / $25.00372K ctx / 128K out
claude-opus-4-7$5.00 / $25.00372K ctx / 128K out
claude-opus-4-8$5.00 / $25.00372K ctx / 128K out
Bestclaude-opus-5-5 (default)$4.00 / $20.00372K ctx / 128K out
claude-opus-5$5.00 / $25.00372K ctx / 128K out
claude-fable-5-1$10.00 / $50.00372K ctx / 128K out
claude-fable-5$10.00 / $50.00372K ctx / 128K out
claude-opus-4-0, claude-opus-4-1$15.00 / $75.00200K ctx / 32K out

Routing defaults: claude-haiku-4-5 (Fast), claude-opus-5-5 (Best)

Run caudra auth login anthropic to sign in to a Claude subscription through browser OAuth. Caudra stores the tokens in its state directory, refreshes them automatically, and shows subscription limits through /usage. Subscription requests always go to api.anthropic.com, even when ANTHROPIC_BASE_URL is set.

This experimental flow uses Claude Code's public client registration. Anthropic limits Pro and Max subscription tokens to official clients in its terms. The flow may stop working when Anthropic changes its OAuth or request protocol.

Recent Claude models accept up to 1M tokens. Caudra runs them at a 372k working window, which keeps cost and latency bounded. That window is an input budget: the model's output allowance sits on top of it rather than inside it, so Caudra holds back less of it before compaction. Add -1m to a model id, like claude-sonnet-4-6-1m, to open the full 1M window instead. Set context_window in providers.toml to pick any other size.

If you already use Claude through AWS Bedrock, you can point Caudra at it instead of the direct Anthropic API. Set CLAUDE_CODE_USE_BEDROCK=1 and Caudra will route all Anthropic requests through Bedrock. The same models, the same features, just a different door.

You will need AWS_REGION and one of the following for auth:

MethodEnv vars
IAM credentialsAWS_ACCESS_KEY_ID + AWS_SECRET_ACCESS_KEY (and optionally AWS_SESSION_TOKEN)
Credentials fileAWS_PROFILE (defaults to default), reads ~/.aws/credentials
Bearer tokenAWS_BEARER_TOKEN_BEDROCK
Gateway proxyCLAUDE_CODE_SKIP_BEDROCK_AUTH=1 + ANTHROPIC_BEDROCK_BASE_URL (skips signing, useful behind a proxy that handles auth)

You can override the model with ANTHROPIC_MODEL and the endpoint with ANTHROPIC_BEDROCK_BASE_URL. These env var names match Claude Code, so if you were already using Bedrock there, the same setup works here.

  • Env var: OPENAI_API_KEY (also supports OAuth device flow)
  • API: https://api.openai.com/v1
MarkerModelsPricing (in/out per 1M tokens)Context
Fastgpt-6-luna (gpt-6 default)$0.10 / $0.50372K ctx / 128K out
Fastgpt-5.6-luna (gpt-5.6 default)$0.20 / $1.20372K ctx / 128K out
Smallgpt-5.4-nano$0.20 / $1.25400K ctx / 128K out
Smallgpt-5.4-mini$0.75 / $4.50400K ctx / 128K out
Smallgpt-4.1-nano$0.10 / $0.401047K ctx / 32K out
Bestgpt-6-astra (gpt-6 default)$10.00 / $50.00372K ctx / 128K out
gpt-6.1-sol$2.00 / $10.00372K ctx / 128K out
gpt-6-sol$2.00 / $10.00372K ctx / 128K out
gpt-5.6-terra$2.00 / $12.00372K ctx / 128K out
Bestgpt-5.6-sol (gpt-5.6 default)$4.00 / $20.00372K ctx / 128K out
gpt-4.1-mini$0.40 / $1.601047K ctx / 32K out
gpt-4.1$2.00 / $8.001047K ctx / 32K out
o4-mini$1.10 / $4.40200K ctx / 100K out
gpt-5.5$5.00 / $30.001050K ctx / 128K out
gpt-5.4$2.50 / $15.001050K ctx / 128K out
o3$2.00 / $8.00200K ctx / 100K out
gpt-5.3-codex$1.75 / $14.00400K ctx / 128K out
gpt-5.2-codex$1.75 / $14.00400K ctx / 128K out
gpt-5.2$1.75 / $14.00400K ctx / 128K out
gpt-5.1-codex-mini$0.25 / $2.00400K ctx / 128K out
gpt-5.1-codex-max$1.25 / $10.00400K ctx / 128K out
gpt-5.1-codex$1.25 / $10.00400K ctx / 128K out

Routing defaults: gpt-6-luna (Fast), gpt-6-astra (Best)

A lane answers inside the release line you are on. On gpt-6 that is gpt-6-luna (Fast) and gpt-6-astra (Best). On gpt-5.6 that is gpt-5.6-luna (Fast) and gpt-5.6-sol (Best).

  • Env var: GEMINI_API_KEY
  • API: https://generativelanguage.googleapis.com/v1beta
  • Features: Native Gemini API with thinking support
MarkerModelsPricing (in/out per 1M tokens)Context
Fastgemini-2.0-flash-lite (default)$0.07 / $0.301048K ctx / 65K out
Bestgemini-2.5-pro (default)$1.25 / $10.001048K ctx / 65K out
gemini-2.5-flash$0.30 / $2.501048K ctx / 65K out

Routing defaults: gemini-2.0-flash-lite (Fast), gemini-2.5-pro (Best)

  • Env var: GH_COPILOT_TOKEN (or run caudra auth login copilot to import a token from gh CLI, the Copilot client, or the system keyring)
  • API: https://api.githubcopilot.com (or GraphQL-discovered Copilot API endpoint)
  • Features: Native Copilot Chat HTTP API with model endpoint discovery
MarkerModelsPricing (in/out per 1M tokens)Context
Smallgpt-5-mini$0.25 / $2.00200K ctx / 100K out
Smallgpt-5.4-mini$0.75 / $4.50200K ctx / 100K out
Smallgpt-5.4-nano$0.20 / $1.25200K ctx / 100K out
Smallclaude-haiku-4.5$1.00 / $5.00200K ctx / 64K out
Smallgemini-3.5-flash$1.50 / $9.00200K ctx / 65K out
Smallmai-code-1-flash-picker$0.75 / $4.50200K ctx / 100K out
Smallgpt-6-luna$0.10 / $0.50200K ctx / 100K out
Fastgpt-5.6-luna (default)$0.20 / $1.20200K ctx / 100K out
gemini-3.6-flash$0.75 / $3.75200K ctx / 65K out
gemini-3.7-flash$0.75 / $3.75200K ctx / 65K out
claude-sonnet-4.5, claude-sonnet-4.6$3.00 / $15.00200K ctx / 64K out
claude-sonnet-5.5$2.00 / $10.00200K ctx / 100K out
claude-sonnet-5$2.00 / $10.00200K ctx / 100K out
gpt-5.5$5.00 / $30.00200K ctx / 100K out
kimi-k2.7-code$0.95 / $4.00200K ctx / 100K out
kimi-k3$3.00 / $15.00200K ctx / 100K out
gemini-3.1-pro-preview$2.00 / $12.00200K ctx / 65K out
gpt-6.1-sol$2.00 / $10.00200K ctx / 100K out
gpt-6-sol$2.00 / $10.00200K ctx / 100K out
gpt-5.4$2.50 / $15.00200K ctx / 100K out
gpt-5.6-sol$4.00 / $20.00200K ctx / 100K out
gpt-5.6-terra$2.00 / $12.00200K ctx / 100K out
gpt-5.3-codex$1.75 / $14.00200K ctx / 100K out
claude-opus-5.5$4.00 / $20.00200K ctx / 128K out
Bestclaude-opus-5, claude-opus-4.8, claude-opus-4.7, claude-opus-4.6, claude-opus-4.5 (default)$5.00 / $25.00200K ctx / 64K out
claude-opus-4.8-fast, claude-fable-5$10.00 / $50.00200K ctx / 100K out
grok-4.5$2.00 / $6.00200K ctx / 100K out
grok-4.6$2.00 / $6.00200K ctx / 100K out

Routing defaults: gpt-5.6-luna (Fast), claude-opus-5 (Best)

  • Env var: OLLAMA_HOST for local/remote (e.g. http://localhost:11434), OLLAMA_API_KEY for auth
  • API: http://localhost:11434/v1
  • Features: Local or remote inference via OLLAMA_HOST, cloud fallback via OLLAMA_API_KEY

This provider talks the OpenAI-compatible /v1 API, so it also works with llama.cpp's server, LocalAI, or anything else that speaks the same protocol. Just point OLLAMA_HOST to the right address (e.g. http://localhost:8080 for llama.cpp).

  • Env var: LLAMA_CPP_API_KEY
  • API: http://localhost:8080/v1
  • Features: Local or remote inference via LLAMA_CPP_HOST, set optional key via LLAMA_CPP_API_KEY

Connects to any OpenAI-compatible /v1 endpoint. Set LLAMA_CPP_HOST to your server address, such as http://localhost:8080, or run caudra auth login llama-cpp, which offers that address and saves your answer to providers.toml. Without either, Caudra reports that LLAMA_CPP_HOST is not set.

  • Env var: MISTRAL_API_KEY
  • API: https://api.mistral.ai/v1
MarkerModelsPricing (in/out per 1M tokens)Context
Fastministral-14b-latest, ministral-14b-2512 (default)$0.20 / $0.20262K ctx
Bestmistral-medium-latest, mistral-medium-3.5, mistral-medium-3-5, mistral-medium-2604 (default)$1.50 / $7.50262K ctx
glm-5-2, zai-glm-5-2$1.40 / $4.401000K ctx
mistral-small-latest, mistral-small-2603$0.15 / $0.60262K ctx

Routing defaults: ministral-14b-latest (Fast), mistral-medium-latest (Best)

  • Env var: ZHIPU_API_KEY (shared across both endpoints)
  • API endpoints:
    • https://api.z.ai/api/paas/v4
    • https://api.z.ai/api/coding/paas/v4
MarkerModelsPricing (in/out per 1M tokens)Context
Fastglm-4.7-flash (default)$0.00 / $0.00200K ctx / 131K out
Smallglm-4.5-flash$0.00 / $0.00131K ctx / 98K out
Smallglm-4.5-air$0.20 / $1.10131K ctx / 98K out
Bestglm-5-code (default)$1.20 / $5.00200K ctx / 131K out
glm-5.2$1.00 / $3.201000K ctx / 131K out
glm-5.1, glm-5$1.00 / $3.20200K ctx / 131K out
glm-4.7, glm-4.6$0.60 / $2.20200K ctx / 131K out
glm-4.5$0.60 / $2.20131K ctx / 98K out

Routing defaults: glm-4.7-flash (Fast), glm-5-code (Best)

  • Env var: DEEPSEEK_API_KEY
  • API: https://api.deepseek.com
  • Features: Thinking mode toggle (on/off), open-weight models
  • Peak pricing: the prices below are off-peak; each turn is billed as it happens, at 2x during 01:00-04:00, 06:00-10:00 UTC
MarkerModelsPricing (in/out per 1M tokens)Context
deepseek-v4-flash$0.22 / $0.661000K ctx / 384K out
Bestdeepseek-v4-pro (default)$0.66 / $1.981000K ctx / 384K out

Routing defaults: deepseek-v4-pro (Best)

  • Env var: OPENROUTER_API_KEY
  • API: https://openrouter.ai/api/v1
  • Features: 300+ models from all providers, prompt caching, provider routing

OpenRouter aggregates models from many providers behind a single API key. Browse available models at openrouter.ai/models. Use any model ID directly (e.g. openrouter/anthropic/claude-sonnet-4).

  • Env var: SYNTHETIC_API_KEY
  • API: https://api.synthetic.new/openai/v1
  • Features: Reasoning effort support (low/medium/high), open-weight models
MarkerModelsPricing (in/out per 1M tokens)Context
Fasthf:zai-org/GLM-4.7-Flash (default)$0.10 / $0.50200K ctx / 131K out
Besthf:moonshotai/Kimi-K2.5 (default)$0.45 / $3.40200K ctx / 131K out
hf:deepseek-ai/DeepSeek-V3.2$0.56 / $1.68200K ctx / 131K out

Routing defaults: hf:zai-org/GLM-4.7-Flash (Fast), hf:moonshotai/Kimi-K2.5 (Best)

  • Env var: TENSORX_API_KEY
  • API: https://api.tensorx.ai/v1
  • Features: Open-weight models, zero data retention, prompt caching

No hardcoded model catalog. Use any model ID supported by this provider.

  • Env var: OPENCODE_API_KEY
  • API: https://opencode.ai/zen/v1
  • Features: Dynamically discovered models via models.dev + all the models provided by Opencode Zen API

No hardcoded model catalog. Use any model ID supported by this provider.

By default Caudra hides free models from the Opencode catalog. To list free models (they use a public fallback, no API key needed), add this to ~/.config/caudra/providers.toml:

[opencode]
enable_free_models = true

The default is false.

  • Env var: XAI_API_KEY (also supports OAuth via caudra auth login xai)
  • API endpoints:
    • https://api.x.ai/v1
    • https://cli-chat-proxy.grok.com/v1
  • Features: OAuth login, account-specific model catalog, Grok reasoning (low/medium/high/xhigh)
MarkerModelsPricing (in/out per 1M tokens)Context
Bestgrok-4.6 (default)$2.00 / $6.00500K ctx / 131K out
grok-4.5$2.00 / $6.00500K ctx / 131K out
grok-4.3$1.25 / $2.501000K ctx / 131K out

Routing defaults: grok-4.6 (Best)

OAuth uses the same first-party xAI client as the official Grok CLI (caudra auth login xai). Browser login (PKCE) is the desktop default; device code is recommended over SSH or in a container. Tokens refresh automatically. After login, Caudra fetches your account catalog from GET /v1/models-v2 on the Grok CLI proxy and caches it for 15 minutes. XAI_BASE_URL only redirects the public API-key endpoint, never the OAuth proxy.

If ~/.grok/auth.json already exists, login offers to reuse it without writing that file.

  • Env var: APERTURE_HOST (e.g. https://your-host.tailnet.ts.net)
  • API: Aperture gateway (set APERTURE_HOST)
  • Features: Tailscale Aperture LLM gateway; set APERTURE_HOST or configure in providers.toml

Aperture discovers models from your gateway. Set APERTURE_HOST to your Tailscale Aperture endpoint (e.g. https://your-host.tailnet.ts.net). No API key needed, Tailscale handles auth.

  • Env var: OPENCODE_API_KEY
  • API: https://opencode.ai/zen/go/v1
  • Features: Dynamically discovered models via models.dev + all the models provided by Opencode Go API

No hardcoded model catalog. Use any model ID supported by this provider. An API key is required.

Models are referenced as provider/model_id:

anthropic/claude-sonnet-4-6
openai/gpt-4.1
xai/grok-4.6
zai/glm-4.7

If the model name is unique across providers, the prefix can be omitted.

providers.toml lives in the config directory (~/.config/caudra/providers.toml on Linux/macOS, %APPDATA%\caudra\providers.toml on Windows). It is the file for provider overrides and custom HTTP providers. Two jobs:

  1. Tweak a built-in (pick a plan, change its base URL, set enable_free_models for Opencode).
  2. Declare a custom provider that speaks OpenAI, Anthropic, or Google wire format.
# Point a built-in at a proxy. Env vars still win over this file.
[anthropic]
base_url = "https://my-proxy.internal"
# Full custom provider. Slug becomes the `provider/` prefix in model specs.
[my-proxy]
display_name = "My Proxy"
protocol = "openai" # openai | openai-responses | anthropic | google
base_url = "https://llm.example.com/v1"
api_key_env = "MY_PROXY_API_KEY"
default_model = "my-proxy/fast-v1"
discover_models = true # also list models via the provider's /models endpoint
[my-proxy.purposes]
fast = "fast-v1" # a prefix: also covers fast-v1-turbo, fast-v1-lora, ...
best = ["smart-v1", "smart-v0"] # a list when one prefix cannot span them
[[my-proxy.models]]
id = "fast-v1"
context_window = 128000
max_output_tokens = 16384
pricing_input = 0.5
pricing_output = 1.5
[[my-proxy.models]]
id = "smart-v1"
context_window = 200000
max_output_tokens = 32000
supports_thinking = true
supports_vision = false

The file can start with version = 1, and a file without it counts as version 1. Caudra writes the key whenever it saves the file. A newer version stops Caudra with an error rather than being misread. Because version belongs to the file, a custom provider cannot use it as a name. See Config file versions.

caudra config example providers prints every providers.toml key with its default, all commented out. Reference configs shows the same text.

FieldTypeDefaultEnvDescription
display_namestringthe built-in name, or the slug-The name pickers and auth status show
protocolstringrequired-The wire format: openai, openai-responses, anthropic, or google
base_urlstringthe plan URL, or the built-in URL<SLUG>_BASE_URLThe API origin. Caudra appends the protocol paths
planstringunset-A built-in plan key, which sets the base URL and the default model
api_key_envstring<SLUG>_API_KEY-The environment variable that holds the API key
api_keystringunset-An API key, stored as plain text. Caudra tries the environment variable and saved credentials first
default_modelstringunset-The model to use after login when none is saved yet, such as my-provider/my-model
discover_modelsboolfalse-Also list the models the provider's model endpoint reports
enable_free_modelsboolunset-Opencode only. Show the free models of its catalog. Unset counts as false
overridestableunset-Aperture only. Overrides for the upstream providers it routes, keyed by upstream id
model_defaultstableunset-Model keys for every model of the provider
purposestableunset-Model id prefixes for the fast and best purposes
modelstable[]unset-The models the provider serves

A built-in slug keeps its compiled protocol, model catalog, and auth setup, so it ignores protocol, api_key_env, discover_models, models and enable_free_models. Opencode still reads enable_free_models.

A [SLUG.purposes] table takes these keys:

FieldTypeDefaultDescription
faststring | string[]unsetModel id prefixes for small, fast models, best first. A prefix covers every id that starts with it, and the first one also names the model that fills the slot, so it has to be a real id
beststring | string[]unsetModel id prefixes for flagship models, best first. A prefix cannot also be in fast. A model a job is bound to in the picker wins over both lists

Each [[SLUG.models]] entry declares one model:

FieldTypeDefaultDescription
idstringrequiredThe model id, which makes the spec SLUG/ID
context_windowintegerdiscovered, or the protocol defaultTokens of context
max_output_tokensintegerdiscovered, or the protocol defaultThe most tokens one response may hold
supports_tool_examplesboolfalseSend tool examples as a structured field. It is off unless declared, because the protocol says nothing about the model behind it
supports_thinkingbooldiscovered, or the protocol defaultThe model accepts extended thinking
requires_thinkingboolfalseFor an API that rejects requests with thinking off. It implies supports_thinking and raises thinking to minimal effort when it is off, compaction included
supports_visionboolfalseThe model accepts images. When false, image input and view_image are off
supports_pdfboolfalseanthropic and openai-responses only. The model reads a PDF that webfetch attaches inside its tool result. When it is off, webfetch returns the text of the PDF instead
supports_cache_breakpointsboolfalseopenai-responses only. The endpoint honours an explicit prompt_cache_breakpoint, so the system prompt closes with one
reasoning_optionstable[]unsetThe reasoning controls the model takes, such as [{ type = "effort", values = ["low", "high"] }]. A type is toggle, effort with values, or budget_tokens with an optional min and max. [] declares that it takes none, so Caudra sends no reasoning level
pricing_inputfloat0USD per million input tokens
pricing_outputfloat0USD per million output tokens
pricing_cache_writefloat0USD per million tokens written to the prompt cache
pricing_cache_readfloat0USD per million tokens read from the prompt cache
pricing_fast_inputfloatunsetUSD per million input tokens in fast mode
pricing_fast_outputfloatunsetUSD per million output tokens in fast mode

A models entry only applies to the exact id it names. When a provider's ids change often, or discover_models finds models you never declared, put the shared settings in model_defaults instead:

[my-proxy.model_defaults]
context_window = 229376
max_output_tokens = 32768
reasoning_options = []
[[my-proxy.models]]
id = "smart-v1"
max_output_tokens = 64000

It takes the same fields as a models entry apart from id, and applies to every model of the provider including discovered ones. A matching models entry wins field by field, so smart-v1 above keeps the 229376-token window and raises only its output cap. Anything a model neither declares nor inherits falls back to discovery, then to the protocol default.

reasoning_options = [] is a declaration, not an omission: it says the endpoint takes no reasoning controls, so Caudra sends no reasoning_effort. Leaving it unset instead lets a thinking level chosen for another model reach an endpoint that rejects it.

A models entry whose id matches no live model is logged once: every setting on it is ignored, which otherwise looks like Caudra disregarding the config.

Custom slugs must not reuse a built-in provider name. A bad TOML parse exits with code 2 at startup so a typo cannot silently empty the registry.

You can also create a custom provider interactively with caudra auth login and choosing the custom option. That writes a starter entry to this file.

Aperture proxies upstream providers, exposing each model as aperture/<upstream>/<model>. Overrides keyed by upstream provider id live under [aperture.overrides]:

[aperture.overrides.llmserver]
base = "llama-cpp"
context_window = 131072
max_output_tokens = 16384
[aperture.overrides.llmserver.models."qwen-3.6"]
context_window = 262144
supports_vision = true

Provider-level fields apply to every model from that upstream. Per-model entries under models win field by field, and take the same keys apart from models. Model ids containing dots must be quoted ("qwen3.6") since TOML treats a bare dotted key as a nested table.

FieldTypeDefaultDescription
context_windowintegerunsetTokens of context
max_output_tokensintegerunsetThe most tokens one response may hold
supports_thinkingboolunsetThe models accept extended thinking
supports_visionboolunsetThe models accept images
basestringunsetThe native provider an opaque upstream works like, such as llama-cpp, google, or anthropic. Caudra warns about a value it does not know and ignores it
path_prefixstring/v1, /v1beta for Gemini routes, none for Anthropic and Z.AIThe path Caudra sends ahead of each request, which Aperture appends to the upstream base URL. Set it to "" when that URL already has its own path
modelstable{}Overrides for single models, keyed by model id, which win key by key. Quote an id that holds a dot, such as models."qwen-3.6"

Caudra sends /v1 (or /v1beta for Gemini routes, nothing for Anthropic and Z.AI), and Aperture appends that path to the upstream's base url. If an upstream base url already carries its own path, set path_prefix = "" for it to avoid a doubled path. Z.AI defaults to no prefix since its API path has no /v1 segment; point the upstream base url at the full API root (e.g. https://api.z.ai/api/paas/v4).

Some built-ins ship multiple plans (different base URLs or default models). caudra auth login <provider> asks which plan to use when more than one exists. You can also set it in TOML:

[mistral]
plan = "coding"
[zai]
plan = "coding"

Current plans:

ProviderPlanWhat it does
MistralstandardStandard at https://api.mistral.ai/v1, default mistral/mistral-medium-latest
MistralcodingVibe / Coding at https://api.mistral.ai/v1, default mistral/mistral-vibe-cli-latest
Z.AIstandardPay-as-you-go at https://api.z.ai/api/paas/v4, default zai/glm-5.1
Z.AIcodingCoding plan at https://api.z.ai/api/coding/paas/v4, default zai/glm-5-code

Env <SLUG>_BASE_URL still wins over both the plan and a base_url in this file.

To add a custom provider or proxy, drop an executable script into the config providers/ directory (~/.config/caudra/providers/ on Linux/macOS, %APPDATA%\caudra\providers\ on Windows). The script must handle these subcommands:

SubcommandTimeoutWhat it does
info5sReturn JSON with display_name, base provider, has_auth
models5sReturn JSON array of model entries (optional)
resolve30sReturn auth JSON (base_url, headers)
logininteractiveOAuth or credential flow
logoutinteractiveClear credentials
refresh30sRefresh auth tokens

resolve is called each time a new agent spawns, so scripts should read tokens from disk instead of caching them in memory. That way auth changes from other processes get picked up.

The base field specifies which built-in provider to inherit the model catalog from. Valid values: anthropic, openai, google, copilot, ollama, llama-cpp, mistral, zai, deepseek, openrouter, synthetic, tensorx, opencode, xai, aperture.

If your provider serves models not in the base catalog, add a models subcommand returning:

[{"id": "my-model-v2", "context_window": 200000, "max_output_tokens": 16384}]

Only id is required. Optional fields: context_window (128K), max_output_tokens (16K), pricing ({input, output, cache_write, cache_read}, all per 1M tokens), supports_tool_examples (defaults to the base provider's setting), supports_thinking (defaults to the base provider's setting), requires_thinking (default false; for APIs that reject requests with thinking off, raises it to minimal effort and implies supports_thinking), supports_vision (defaults to the base provider's setting; when false, image input and the view_image tool are disabled). Without this subcommand, the base provider's models are used.

A llama-cpp model can replace Caudra's token-budget mapping with its native thinking fields. Each thinking mode maps to a JSON fragment merged into the request body:

[{
"id": "reasoning-model",
"supports_thinking": true,
"thinking_fields": {
"off": {"reasoning_effort": "none"},
"adaptive": {"reasoning_effort": "medium"},
"low": {"reasoning_effort": "low"},
"medium": {"reasoning_effort": "medium"},
"xhigh": {"reasoning_effort": "xhigh"}
}
}]

off is used when thinking is off, adaptive when thinking is on without a chosen level. Any other key is an effort level, one of none, minimal, low, medium, high, xhigh, max. The levels you declare are the ones the model accepts: whatever you ask for snaps into them, downwards first, so a level the model never advertised is never sent. Every part is optional.

Fragments are merged into the body, so nesting works too. A template toggle is just a fragment:

"thinking_fields": {
"off": {"chat_template_kwargs": {"enable_thinking": false}},
"adaptive": {"chat_template_kwargs": {"enable_thinking": true}}
}

Named modes send only these fields, no token budget. An explicit /thinking <budget> snaps into the levels you declared; a model that declares none gets the adaptive fragment plus thinking_budget_tokens. Any mode you left undeclared falls back to the usual thinking_budget_tokens mapping, so no request ever ends up saying nothing. Models without thinking_fields keep the existing llama.cpp behavior.

Dynamic provider models are namespaced as {slug}/{model_id} (e.g. myproxy/claude-sonnet-4-6).

  • Must start with a letter or digit
  • Only letters, digits, underscores, and hyphens after that
  • Can't reuse a built-in provider's slug
  • Must be executable

Website privacy and analytics