Skip to content
This documentation covers the kagent 1.0 alpha. For the latest 0.x release, see the 0.x docs.

For the complete documentation index, see llms.txt. Markdown versions of all docs pages are available by appending .md to any docs URL.

Anthropic

Page as Markdown

Configure kagent to use Anthropic Claude models by creating a ModelConfig for the Anthropic provider.

The Anthropic provider calls the Anthropic API directly.

Note

This provider works on the kagent, byo, and claude runtimes, but not on codex. A claude Harness accepts no anthropic settings beyond baseUrl. For more information, see Agent harness. Claude models served through Google Cloud Vertex AI do not run on kagent 1.0. For the reason and the alternatives, see Google Vertex AI.

Create the ModelConfig

  1. Save your Anthropic API key as an environment variable.

    export ANTHROPIC_API_KEY=<your_api_key>
  2. Create a Kubernetes Secret that stores the API key. Create it in the same namespace as the AgentTemplates that use it, such as kagent.

    kubectl create secret generic kagent-anthropic -n kagent --from-literal ANTHROPIC_API_KEY=$ANTHROPIC_API_KEY
  3. Create a ModelConfig that references the Secret.

    kubectl apply -f - <<EOF
    apiVersion: api.kagent.dev/v1alpha3
    kind: ModelConfig
    metadata:
      name: anthropic-model-config
      namespace: kagent
    spec:
      apiKeySecret: kagent-anthropic
      apiKeySecretKey: ANTHROPIC_API_KEY
      model: claude-sonnet-5
      provider: Anthropic
      anthropic: {}
    EOF
    FieldDescription
    apiKeySecretThe name of the Kubernetes Secret that stores the API key, in the same namespace as this ModelConfig.
    apiKeySecretKeyThe key within that Secret that holds the API key.
    modelThe model to use. For the available models, see the Anthropic model docs.
    providerThe provider to use, Anthropic.
    anthropicSettings that only the Anthropic provider takes. An empty block is valid.

Anthropic provider settings

The anthropic block takes the following optional settings. For every field, including its type, default, and validation rules, see the API reference.

FieldDescription
baseUrlAn alternative API endpoint, for a proxy or a compatible service.
maxTokensA cap on the tokens generated in one response.
temperatureHow much randomness the model applies when it picks the next token.
topPThe nucleus sampling cutoff.
topKHow many candidate tokens to sample from.
promptCachingWhether to bill the reusable prefix of each request as a cache read instead of fresh input. Defaults to false. For more information, see Prompt caching.
cacheTTLHow long Anthropic retains a cached prefix. Applies only when promptCaching is true. Supported values are 5m (default) or 1h.

Prompt caching

An agent that calls a model many times for one task resends the same prefix every time, including the tool definitions, the system prompt, and the turns already taken. When promptCaching is true, kagent marks that prefix with cache_control breakpoints, and Anthropic bills a later request that reuses it at a fraction of the normal input price. Because the conversation breakpoint moves with every turn, each call in an agent loop reads the whole previous history from the cache and writes only the new turn.

Enable the field wherever a tool-using agent makes many model calls per task against a stable system prompt and tool set. Without it, the full history is billed as fresh input on every call.

spec:
  anthropic:
    promptCaching: true
    cacheTTL: "5m"

Two properties of Anthropic’s pricing decide whether caching reduces cost. Neither one produces an error or a warning when caching ends up costing more.

  • A cache write costs more than ordinary input. A prefix must be read at least once before the saving on reads exceeds the premium charged on the write, so caching a prompt that is used once costs more than not caching it.
  • Each model sets a minimum cacheable prefix, between 1024 and 4096 tokens depending on the model. Under that minimum Anthropic ignores the breakpoints silently, so the request succeeds, the response is normal, and nothing is cached.

Important

1h is not an improvement on 5m. Anthropic bills 1-hour cache writes at a higher per-token rate than 5-minute writes, and every cache hit refreshes the window, so an agent loop whose calls are less than 5 minutes apart keeps its prefix cached for the whole task on 5m. Choose 1h only when a task’s model calls are spaced far enough apart that a 5-minute cache would expire between them. Otherwise, the higher write rate adds cost without reducing it anywhere else.

For the models that support caching, the current minimum prefix sizes, and the pricing, see the Anthropic prompt caching docs.

Warning

The claude harness rejects promptCaching: true. That harness accepts no anthropic settings beyond baseUrl, so a ModelConfig that enables caching fails to compile for it, and the AgentTemplate reports Claude does not support Anthropic provider options beyond baseUrl yet rather than becoming ready. Claude Code caches its own prefix on a 5-minute window regardless, so the setting gains nothing there. Where a claude agent and a kagent agent must share one installation, give the claude agent a ModelConfig that leaves promptCaching unset rather than enabling the field chart-wide. For the settings that each harness takes, see Agent harness.

Prompt caching at install time

Setting providers.anthropic.config in the Helm chart writes these fields into the ModelConfig that the chart generates, which saves editing that resource after every install.

providers:
  anthropic:
    config:
      promptCaching: true
      cacheTTL: "5m"

The setting reaches only the generated ModelConfig. A ModelConfig that you create yourself, including the one in Create the ModelConfig, takes the fields in its own spec.anthropic block.

Use the ModelConfig

Reference the ModelConfig by name from an AgentTemplate in the same namespace.

spec:
  modelConfig:
    name: anthropic-model-config

Next steps