Skip to content
This documentation covers the kagent 1.0 alpha. For the latest 0.x release, see the 0.x docs.

For the complete documentation index, see llms.txt. Markdown versions of all docs pages are available by appending .md to any docs URL.

Tracing

Page as Markdown

Enable OpenTelemetry tracing for kagent, then read a trace that runs from the controller through to the Actor that executed your agent.

A trace records one agent request as a tree of timed spans, so you can see where a slow or failed request spent its time and which model and tool calls it made along the way. In kagent 1.0 a single request crosses the controller, the Agent Substrate router, and the ActorActorThe sandboxed unit of compute, provided by Agent Substrate, that runs a Session's conversation loop. Every Session is backed by one.Learn more that runs the agent, and one trace ties all three together.

About trace coverage

A trace follows a W3C Trace Context header that the controller passes along with each request. The following diagram traces one request from the caller to the agent runtime.

    flowchart LR
    caller["Caller"]
    subgraph controllerproc["kagent controller"]
        grpc["gRPC API"]
        gateway["A2A gateway"]
    end
    subgraph substrateproc["Agent Substrate"]
        router["Router"]
    end
    subgraph actorproc["Actor"]
        runtime["Agent runtime"]
    end
    %% Cross-subgraph edges are declared outside every subgraph block, because
    %% mermaid assigns a node to the subgraph that first references it.
    caller --> grpc
    grpc --> gateway
    gateway -->|traceparent| router
    router --> runtime
    classDef boundary fill:#a78bfa26,stroke:#a78bfa,stroke-width:2px
    classDef inner fill:#80808033,stroke:#9ca3af,stroke-width:1px
    class controllerproc,substrateproc,actorproc boundary
    class grpc,gateway,router,runtime inner
  

A caller reaches the gRPC API on the kagent controller, which starts the trace. The controller hands the request to its A2A gateway, which opens an A2AA2AThe Agent-to-Agent protocol, which callers and other agents use to talk to an Agent. The conversation's context identifier is the Session ID, so a second message on the same ID continues the same conversation.Learn more (Agent-to-Agent) connection to the Session’s Actor and injects a traceparent header into that call. The Agent Substrate router forwards the call to the Worker that runs the Actor, and adds its own spans to the trace. The agent runtime inside the Actor reads the header and continues the same trace, so the model and tool spans it produces hang off the controller’s spans rather than starting a trace of their own.

Important

The controller passes its tracing configuration to the kagent, codex, and claude runtimes. Each of the three exports on its own instrumentation, so the span names in this page describe the kagent runtime and do not carry over to the other two. An agent on the byo runtime receives no tracing configuration, and its half of the trace is missing. For the available runtimes, see Choose a runtime.

Note

A byo image that implements OTel itself reads the exporter variables from the Harness spec.env, which the controller leaves alone for this runtime. Its spans still do not reach a collector inside the cluster, because kagent adds the collector to an Actor’s egress allowlist only for the runtimes it configures, and no field adds a host to that list by hand. For more information, see Networking and egress control.

Each hop reports itself as a separate OpenTelemetry (OTel) service. A tracing backend uses these service names to group the spans.

  • The controller reports as kagent-controller in the kagent service namespace. Its spans also carry the pod, node, and namespace that the controller runs on.
  • The Agent Substrate router reports as two services, because its pod runs two containers. The router’s own spans, such as its lookup of the Actor for a request, report as atenet-router. The spans of the agentgateway proxy that forwards the request to the Worker report as agentgateway. Only the agentgateway spans join the agent request trace. The atenet-router spans form separate Agent Substrate traces.
  • Each agent runtime reports as its own service, named for the AgentAgentA Kubernetes custom resource that pairs one AgentTemplate with one Harness. Each side takes either an inline spec or a reference to an existing resource, and the controller compiles the pair into a revision.Learn more it was compiled from. The my-first-agent Agent reports as my-first-agent.

Note

A service per Agent is a change from kagent 0.x, where every agent reported under one kagent service. A backend that you filter by service now shows one entry for each Agent, and adding an Agent adds a service.

Spans

The kagent runtime creates the same spans for every agent, and most span names describe the operation rather than the agent. The invoke_agent span is the exception, because its name carries the name of the agent that ran. To narrow a search to one agent, filter by service name rather than by span name. The following spans appear in nesting order, from the span that accepts the request down to the model and tool calls that serve it.

SpanWhen it is created
lf.a2a.v1.A2AService/SendMessageOnce per request, as the root of the runtime’s half of the trace. The runtime creates it when it accepts the A2A call from the controller. The controller reports spans of the same name for its own side of the call.
a2a.requestOnce per request. Records the A2A method and the final state of the task in the a2a.method and a2a.task.state attributes.
invoke_agent <agent>Once per request, named for the AgentAgentA Kubernetes custom resource that pairs one AgentTemplate with one Harness. Each side takes either an inline spec or a reference to an existing resource, and the controller compiles the pair into a revision.Learn more that serves it, with each hyphen replaced by an underscore. The Agent my-first-agent produces invoke_agent my_first_agent, while its service name keeps the hyphens, so the two spellings differ.
generate_content <model>Once per model call, named for the model that was called.
execute_tool <tool>Once per tool call, named for the tool that was called. Records the call’s arguments and the tool’s reply in the gcp.vertex.agent.tool_call_args and gcp.vertex.agent.tool_response attributes.
execute_tool (merged)Once per model turn that calls more than one tool, as the parent of that turn’s execute_tool spans. A turn that calls a single tool creates no merged span.

Correlation attributes

A trace tells you which request you are looking at through attributes on its spans, not through the span names. The runtime stamps the following four attributes onto its root span and copies them onto every descendant span. A search on any one of these attributes returns the whole subtree rather than a single span.

AttributeValue
a2a.task.idThe A2A task ID, which identifies one turn of a conversation.
gen_ai.conversation.idThe A2A context ID, which identifies the conversation and is stable across its turns.
gen_ai.agent.idThe Agent, as <namespace>/<name>. The shorter gen_ai.agent.name carries the name alone.
enduser.idThe authenticated caller, such as admin@kagent.dev.
kagent.runtimeThe runtime that served the request, such as adk-go.

The runtime also adds each scalar value in the A2A message’s metadata as an a2a.message.metadata.<key> attribute, so a client can tag a request and search for it later. Unlike the correlation attributes, these tags stay on the a2a.request span alone, so a search on one returns that span instead of the whole subtree.

Warning

When the otel.capture.messageContent Helm setting is true, prompts and replies reach your tracing backend. On the kagent runtime, the generate_content span of each model call then carries the conversation as the gen_ai.input.messages and gen_ai.output.messages attributes, and the Agent’s system prompt as gen_ai.system_instructions. The setting defaults to false, which omits all three.

The setting does not govern tool content. An execute_tool span carries the call’s arguments and the tool’s reply in gcp.vertex.agent.tool_call_args and gcp.vertex.agent.tool_response whether the setting is true or false, so a tool that returns sensitive data sends it to your tracing backend on the default settings. Turn tracing off for an agent whose tools return data that must not leave the cluster. For how to use this content as an audit record, see Audit prompts.

Before you begin

  1. Install kagent.
  2. Create your first agent, so that you have an AgentAgentA Kubernetes custom resource that pairs one AgentTemplate with one Harness. Each side takes either an inline spec or a reference to an existing resource, and the controller compiles the pair into a revision.Learn more to send a request to. That guide also installs the kagent CLI. The steps on this page need the 1.0.0-alpha8 CLI, because the CLIs of other releases, newer ones included, do not have the Session commands that these steps use. To check your version, run kagent version.
  3. Set up a tracing backend. The OTel stack sends traces to Tempo, and the Lightweight OTel stack sends traces to Jaeger. Both guides turn on tracing for you, so you can skip to Review a trace.

Enable tracing

Tracing is off by default. Turning it on is a Helm change, because the controller reads its tracing configuration from the environment and passes that configuration to the agent runtimes that the controller starts. The following steps send traces to the collector that both stack guides install. To send traces to another OTLP backend, change the endpoint.

Important

The telemetry values changed shape in 1.0. otel.tracing.*, otel.logging.*, otel.captureSensitiveContent, and the insecure flag were removed, and otel.exporter.otlp.*, otel.traces.*, otel.logs.*, and otel.capture.* replace them. The chart defines no check that rejects a removed setting, so neither Helm flag that reuses stored values carries you across. --reuse-values restores the previous chart’s defaults, which leaves otel.traces and otel.metrics undefined and fails the upgrade while the controller ConfigMap renders. --reset-then-reuse-values completes, but the chart ignores the removed keys, so a release that set otel.tracing.enabled to true turns tracing off without reporting an error. If your release still holds the old settings, write the replacements into a values file and upgrade with --values, as shown in the following steps. On subsequent upgrades, --reuse-values works again.

  1. Save the current revision of your Agent. A later step uses it to tell when kagent recompiles the Agent with the new settings. The command first waits for any recompile that is still in progress, such as one from an earlier Helm upgrade, so that it saves a finished revision.

    for i in $(seq 1 60); do
      REVISIONS=$(kubectl get agent my-first-agent -n kagent \
        -o jsonpath='{.status.desiredRevision} {.status.latestSuccessfulRevision}')
      [ "${REVISIONS% *}" = "${REVISIONS#* }" ] && break
      sleep 5
    done
    export OLD_REVISION=${REVISIONS#* }
    echo "Current revision: $OLD_REVISION"
  2. Get your current Helm values for kagent.

    helm get values kagent -n kagent -o yaml > values.yaml
  3. Add the tracing settings to the values file.

    otel:
      exporter:
        otlp:
          endpoint: http://otel-collector.telemetry.svc.cluster.local:4317
          protocol: grpc
          timeout: 15000
      traces:
        enabled: true
    Review the following table to understand this configuration.
    FieldDescription
    traces.enabledWhether to export traces at all. Defaults to false.
    exporter.otlp.endpointThe OTLP endpoint that every signal exports to, as an http:// or https:// URL. An http:// endpoint sends plaintext. Empty by default. When a signal is enabled and neither this setting nor its per-signal counterpart holds an endpoint, the controller disables that signal and reports OTLP traces endpoint is required when traces export is enabled.
    exporter.otlp.protocolgrpc or http/protobuf. Defaults to grpc, which matches the port 4317 in the example endpoint. Point http/protobuf at port 4318 instead.
    exporter.otlp.timeoutThe export timeout in milliseconds. Empty by default, which keeps the OTel SDK default.
    traces.endpoint, traces.protocolSend traces somewhere other than the other signals. Each one overrides its exporter.otlp counterpart for traces alone. Both are empty by default.

    An endpoint is an absolute http:// or https:// URL, and one that carries credentials, a query, or a fragment leaves its signal disabled in the same way. An invalid telemetry setting never fails the upgrade or the Harness. The controller reports each one as a warning when it starts, as the log record invalid agent telemetry configuration; disabling signal, and turns off only the signal that the setting belongs to. Check that log after you change these values, because no other surface reports the problem and an agent with a disabled signal runs normally.

    The two endpoint settings differ in how the controller treats the path. A per-signal endpoint such as traces.endpoint is used exactly as you write it. The shared exporter.otlp.endpoint is used as written on the grpc protocol, and gains a /v1/traces suffix on http/protobuf, so point the shared setting at the collector’s root rather than at a signal path.

  4. Upgrade the kagent Helm release.

    helm upgrade kagent \
      oci://ghcr.io/kagent-dev/kagent/helm/kagent \
      --version 1.0.0-alpha8 \
      --namespace kagent \
      --values values.yaml
  5. Wait for kagent to recompile the Agent. The controller rebuilds each Agent after the controller restarts. A Session that you create before the rebuild finishes starts from the previous revision, without the new settings. The following command prints Recompiled when the new revision is ready.

    for i in $(seq 1 60); do
      [ "$(kubectl get agent my-first-agent -n kagent \
        -o jsonpath='{.status.latestSuccessfulRevision}')" != "$OLD_REVISION" ] \
        && echo "Recompiled" && break
      sleep 5
    done

    If the command finishes without printing Recompiled, the upgrade did not change the settings that kagent compiles into the Agent. Either the settings were already in place, or the chart did not recognize the otel keys. Helm accepts a key that a chart does not define without an error, so check that you upgraded to version 1.0.0-alpha8 of the chart, which uses the keys on this page.

  6. Create a new Session, so that its Actor starts from a runtime that has the tracing configuration.

    kagent agent session create --agent my-first-agent
  7. Confirm that the Session runs the Agent’s current revision. The command waits until kagent finishes compiling the Agent, then compares that revision with the one that the Session started from. If the command prints Outdated, the Session was created from an earlier revision, and exports without the new settings. Create another Session, and run the command again.

    for i in $(seq 1 60); do
      REVISIONS=$(kubectl get agent my-first-agent -n kagent \
        -o jsonpath='{.status.desiredRevision} {.status.latestSuccessfulRevision}')
      [ "${REVISIONS% *}" = "${REVISIONS#* }" ] && break
      sleep 5
    done
    SESSION_REVISION=$(kagent agent session get $SESSION_ID -o json | jq -r '.session.preparedRevision')
    [ "$SESSION_REVISION" = "${REVISIONS#* }" ] && echo "Current" || echo "Outdated"

Review a trace

Send a request to a new Session, then find its trace in the backend that you set up.

  1. Send a request to a new Session to produce a trace.

    export SESSION_ID=$(kagent agent session create --agent my-first-agent -o json | jq -r '.session.id')
    kagent agent invoke --session $SESSION_ID --task "What is 2+2?"
  2. Open the trace in your tracing backend.

    1. Forward the Grafana port, and leave the command running.
      kubectl port-forward -n telemetry svc/kube-prometheus-stack-grafana 3000:80
    2. In your browser, open Grafana at http://localhost:3000, and log in. For the password, see Explore the telemetry in Grafana.
    3. Open Explore, select the Tempo data source, and select the Search query type.
    4. From the Service Name list, select my-first-agent, and run the query. Selecting kagent-controller instead returns the same traces from the controller’s side.
    5. Click a trace to open it.

  3. Review the span tree. The trace starts with the controller’s spans, continues through the Agent Substrate router, and ends with the agent runtime’s spans. The following example shows the spans of one request, with the service that reported each span.

    lf.a2a.v1.A2AService/SendMessage                         kagent-controller
      lf.a2a.v1.A2AService/SendMessage                       kagent-controller
        POST /*                                              agentgateway
          POST                                               agentgateway
            lf.a2a.v1.A2AService/SendMessage                 my-first-agent
              a2a.request                                    my-first-agent
                invoke_agent my_first_agent                   my-first-agent
                  generate_content gpt-4.1-mini              my-first-agent
                    HTTP POST                                my-first-agent
    
  4. To narrow a search to one conversation, search by a correlation attribute, such as gen_ai.conversation.id=<context-id>.

Agent Substrate traces

Agent Substrate records traces for its own work, such as scheduling an Actor onto a Worker and restoring it from a snapshot. These traces are separate from the agent request trace. Agent Substrate traces do not share the request trace’s ID, so a request trace does not show how long the Actor took to resume. To investigate a slow start, look up the Agent Substrate traces from the same time window, or read the Actor’s suspend and resume records, which carry the trace ID of each operation.

ServiceReports
atenet-routerRequests that the router receives, and its calls to ateapi to find or resume the Actor for each request.
ateapiActor lifecycle operations, such as create, resume, and suspend, and the scheduling of Actors onto Workers.
ateletWork on a Worker’s node, such as restoring an Actor from a snapshot.
ateom-gvisorWork inside the sandbox that runs the Actor.
atecontrollerReconciliation of Agent Substrate resources, such as WorkerPools.

Agent Substrate exports traces only when its Helm release sets otel.endpoint, and it keeps 1% of its traces by default. To keep more, raise otel.traces.samplingRatio, as the stack guides do. For the steps, see Send Agent Substrate telemetry to the collector.

Traces from a suspended Actor

Agent Substrate checkpointsCheckpointA durable pin on the snapshot that a Session most recently suspended to, and a record of how far its transcript had advanced. Not a new state: tagging copies the snapshot so that Agent Substrate does not collect it, and a second Session can be forked from it.Learn more an Actor as soon as the response body closes, which is sooner than a batching span exporter normally sends its buffer. Spans still in the buffer at that moment freeze inside the snapshotSnapshotThe stored state that an Actor suspends to, held in object storage. Resuming restores the Actor from its most recent snapshot, which is what makes suspending idle agents cheap.Learn more and reach the backend only when the session next resumes, or never at all for a conversation’s last message.

To avoid losing them, the kagent, codex, and claude runtimes flush their span buffer after each A2A handler returns, before the response completes. The flush is unconditional and waits up to three seconds, and no setting changes either. An agent on the byo runtime flushes only if its own image does, so a conversation’s last turn can lose its spans there.

The flush lets a kagent trace arrive promptly rather than on the exporter’s own schedule. To understand what suspension does to an Actor, see Suspend and resume.

Turn tracing off

Turn off the trace exporter, then create a new Session so that the change takes effect.

  1. Disable tracing in the kagent Helm release.

    helm upgrade kagent \
      oci://ghcr.io/kagent-dev/kagent/helm/kagent \
      --version 1.0.0-alpha8 \
      --namespace kagent --reuse-values \
      --set otel.traces.enabled=false

    Note

    Turning tracing off compiles an explicit off state rather than an absent one. The controller sets the trace exporter to none in each runtime, so an agent never falls back to the OpenTelemetry SDK’s own default of exporting to localhost. The controller sets OTEL_SDK_DISABLED to true only when traces, metrics, and logs are all off. Both stack guides turn logs on, so the SDK stays enabled after this step.

  2. Create a new Session to pick up the change, because an existing Actor keeps the configuration it started with.

  3. To remove the tracing backend, follow the cleanup steps in the OTel stack or Lightweight OTel stack guide.

Next steps