Skip to content
This documentation covers the kagent 1.0 alpha. For the latest 0.x release, see the 0.x docs.

For the complete documentation index, see llms.txt. Markdown versions of all docs pages are available by appending .md to any docs URL.

Tracing

Page as Markdown

Enable OpenTelemetry tracing for kagent, then read a trace that runs from the controller through to the Actor that executed your agent.

A trace records one agent request as a tree of timed spans, so you can see where a slow or failed request spent its time and which model and tool calls it made along the way. In kagent 1.0 a single request crosses the controller, the Agent Substrate router, and the ActorActorThe sandboxed unit of compute, provided by Agent Substrate, that runs a Session's conversation loop. Every Session is backed by one.Learn more that runs the agent, and one trace ties all three together.

About trace coverage

A trace follows a W3C Trace Context header that the controller passes along with each request. The following diagram traces one request from the caller to the agent runtime.

    flowchart LR
    caller["Caller"]
    subgraph controllerproc["kagent controller"]
        grpc["gRPC API"]
        gateway["A2A gateway"]
    end
    subgraph substrateproc["Agent Substrate"]
        router["Router"]
    end
    subgraph actorproc["Actor"]
        runtime["Agent runtime"]
    end
    %% Cross-subgraph edges are declared outside every subgraph block, because
    %% mermaid assigns a node to the subgraph that first references it.
    caller --> grpc
    grpc --> gateway
    gateway -->|traceparent| router
    router --> runtime
    classDef boundary fill:#a78bfa26,stroke:#a78bfa,stroke-width:2px
    classDef inner fill:#80808033,stroke:#9ca3af,stroke-width:1px
    class controllerproc,substrateproc,actorproc boundary
    class grpc,gateway,router,runtime inner
  

A caller reaches the gRPC API on the kagent controller, which starts the trace. The controller hands the request to its A2A gateway, which opens an A2AA2AThe Agent-to-Agent protocol, which callers and other agents use to talk to an Agent. The conversation's context identifier is the Session ID, so a second message on the same ID continues the same conversation.Learn more (Agent-to-Agent) connection to the Session’s Actor and injects a traceparent header into that call. The Agent Substrate router forwards the call to the Worker that runs the Actor, and adds its own spans to the trace. The agent runtime inside the Actor reads the header and continues the same trace, so the model and tool spans it produces hang off the controller’s spans rather than starting a trace of their own.

Important

The controller passes its tracing configuration to the kagent, codex, and claude runtimes. Each of the three exports on its own instrumentation, so the span names in this page describe the kagent runtime and do not carry over to the other two. An agent on the byo runtime receives no tracing configuration, and its half of the trace is missing. For the available runtimes, see Choose a runtime.

Note

A byo image that implements OTel itself reads the exporter variables from the Harness spec.env, which the controller leaves alone for this runtime. Its spans still do not reach a collector inside the cluster, because kagent adds the collector to an Actor’s egress allowlist only for the runtimes it configures, and no field adds a host to that list by hand. For more information, see Networking and egress control.

Each hop reports itself as a separate OpenTelemetry (OTel) service. A tracing backend uses these service names to group the spans.

  • The controller reports as kagent-controller in the kagent service namespace. Its spans also carry the pod, node, and namespace that the controller runs on.
  • The Agent Substrate router reports as two services, because its pod runs two containers. The router’s own spans, such as its lookup of the Actor for a request, report as atenet-router. The spans of the agentgateway proxy that forwards the request to the Worker report as agentgateway. Only the agentgateway spans join the agent request trace. The atenet-router spans form separate Agent Substrate traces.
  • Each agent runtime reports as its own service, named for the AgentAgentA Kubernetes custom resource that pairs one AgentTemplate with one Harness. Each side takes either an inline spec or a reference to an existing resource, and the controller compiles the pair into a revision.Learn more it was compiled from. The my-first-agent Agent reports as my-first-agent.

Note

A service per Agent is a change from kagent 0.x, where every agent reported under one kagent service. A backend that you filter by service now shows one entry for each Agent, and adding an Agent adds a service.

Spans

The kagent runtime creates the same spans for every agent, and most span names describe the operation rather than the agent. The invoke_agent span is the exception, because its name carries the service name of the agent that ran. To narrow a search to one agent, filter by service name rather than by span name. The following spans appear in nesting order, from the span that accepts the request down to the model and tool calls that serve it.

SpanWhen it is created
POST /lf.a2a.v1.A2AService/SendMessageOnce per request, as the root of the runtime’s half of the trace. The runtime creates it when it accepts the A2A call from the controller.
a2a.requestOnce per request. Records the A2A method and the final state of the task in the a2a.method and a2a.task.state attributes.
invocationOnce per request, as the parent of the agent’s own work.
invoke_agent <agent>Once per request, named for the AgentAgentA Kubernetes custom resource that pairs one AgentTemplate with one Harness. Each side takes either an inline spec or a reference to an existing resource, and the controller compiles the pair into a revision.Learn more that serves it, such as invoke_agent my-first-agent. The name matches the runtime’s service name.
generate_content <model>Once per model call, named for the model that was called.
execute_tool <tool>Once per tool call, named for the tool that was called.
execute_tool (merged)Once per model turn that calls more than one tool, as the parent of that turn’s execute_tool spans. A turn that calls a single tool creates no merged span.

Correlation attributes

A trace tells you which request you are looking at through attributes on its spans, not through the span names. The runtime stamps the following four attributes onto its root span and copies them onto every descendant span. A search on any one of these attributes returns the whole subtree rather than a single span.

AttributeValue
gen_ai.task.idThe A2A task ID, which identifies one turn of a conversation.
gen_ai.conversation.idThe A2A context ID, which identifies the conversation and is stable across its turns.
kagent.app_nameThe AgentTemplate, as <namespace>__NS__<name> with hyphens replaced by underscores.
kagent.user_idThe authenticated caller, or A2A_USER_<context-id> for an unauthenticated one.

The runtime also adds each scalar value in the A2A message’s metadata as an a2a.message.metadata.<key> attribute, so a client can tag a request and search for it later. Unlike the four correlation attributes, these tags stay on the invocation span alone, so a search on one returns that span instead of the whole subtree.

Warning

When the otel.captureSensitiveContent Helm setting is true, prompts and replies reach your tracing backend. The spans for a model call then carry the full serialized request and response as the gcp.vertex.agent.llm_request and gcp.vertex.agent.llm_response attributes, truncated to a prefix when a payload is larger than 32 KiB. The setting defaults to false, which leaves both attributes as {}. For how to use this content as an audit record, see Audit prompts.

Before you begin

  1. Install kagent.
  2. Create your first agent, so that you have an AgentAgentA Kubernetes custom resource that pairs one AgentTemplate with one Harness. Each side takes either an inline spec or a reference to an existing resource, and the controller compiles the pair into a revision.Learn more to send a request to. That guide also installs the kagent CLI. The steps on this page need the 1.0.0-alpha7 CLI, because the CLIs of other releases, newer ones included, do not have the Session commands that these steps use. To check your version, run kagent version.
  3. Set up a tracing backend. The OTel stack sends traces to Tempo, and the Lightweight OTel stack sends traces to Jaeger. Both guides turn on tracing for you, so you can skip to Review a trace.

Enable tracing

Tracing is off by default. Turning it on is a Helm change, because the controller reads its tracing configuration from the environment and passes that configuration to the agent runtimes that the controller starts. The following steps send traces to the collector that both stack guides install. To send traces to another OTLP backend, change the endpoint.

  1. Save the current revision of your Agent. A later step uses it to tell when kagent recompiles the Agent with the new settings. The command first waits for any recompile that is still in progress, such as one from an earlier Helm upgrade, so that it saves a finished revision.

    for i in $(seq 1 60); do
      REVISIONS=$(kubectl get agent my-first-agent -n kagent \
        -o jsonpath='{.status.desiredRevision} {.status.latestSuccessfulRevision}')
      [ "${REVISIONS% *}" = "${REVISIONS#* }" ] && break
      sleep 5
    done
    export OLD_REVISION=${REVISIONS#* }
    echo "Current revision: $OLD_REVISION"
  2. Get your current Helm values for kagent.

    helm get values kagent -n kagent -o yaml > values.yaml
  3. Add the tracing settings to the values file.

    otel:
      tracing:
        enabled: true
        exporter:
          otlp:
            endpoint: http://otel-collector.telemetry.svc.cluster.local:4317
            protocol: grpc
            timeout: 15000
            insecure: true
    Review the following table to understand this configuration.
    FieldDescription
    enabledWhether to export traces at all. Defaults to false.
    exporter.otlp.endpointThe OTLP endpoint to export to. Empty by default, which leaves the exporter on the OTel default of localhost:4317.
    exporter.otlp.protocolgrpc or http/protobuf. Defaults to grpc, which matches the port 4317 in the example endpoint. Point http/protobuf at port 4318 instead.
    exporter.otlp.timeoutThe export timeout in milliseconds. Defaults to 15000.
    exporter.otlp.insecureWhether to skip Transport Layer Security (TLS) for the exporter connection. Defaults to true.
  4. Upgrade the kagent Helm release.

    helm upgrade kagent \
      oci://ghcr.io/kagent-dev/kagent/helm/kagent \
      --version 1.0.0-alpha7 \
      --namespace kagent \
      --values values.yaml
  5. Wait for kagent to recompile the Agent. The controller rebuilds each Agent after the controller restarts. A Session that you create before the rebuild finishes starts from the previous revision, without the new settings. The following command prints Recompiled when the new revision is ready.

    for i in $(seq 1 60); do
      [ "$(kubectl get agent my-first-agent -n kagent \
        -o jsonpath='{.status.latestSuccessfulRevision}')" != "$OLD_REVISION" ] \
        && echo "Recompiled" && break
      sleep 5
    done

    If the command finishes without printing Recompiled, the upgrade did not change the settings that kagent compiles into the Agent. Either the settings were already in place, or the chart did not recognize the otel keys. Helm accepts a key that a chart does not define without an error, so check that you upgraded to version 1.0.0-alpha7 of the chart, which uses the keys on this page.

  6. Create a new Session, so that its Actor starts from a runtime that has the tracing configuration.

    kagent agent session create --agent my-first-agent
  7. Confirm that the Session runs the Agent’s current revision. The command waits until kagent finishes compiling the Agent, then compares that revision with the one that the Session started from. If the command prints Outdated, the Session was created from an earlier revision, and exports without the new settings. Create another Session, and run the command again.

    for i in $(seq 1 60); do
      REVISIONS=$(kubectl get agent my-first-agent -n kagent \
        -o jsonpath='{.status.desiredRevision} {.status.latestSuccessfulRevision}')
      [ "${REVISIONS% *}" = "${REVISIONS#* }" ] && break
      sleep 5
    done
    SESSION_REVISION=$(kagent agent session get $SESSION_ID -o json | jq -r '.session.preparedRevision')
    [ "$SESSION_REVISION" = "${REVISIONS#* }" ] && echo "Current" || echo "Outdated"

Review a trace

Send a request to a new Session, then find its trace in the backend that you set up.

  1. Send a request to a new Session to produce a trace.

    export SESSION_ID=$(kagent agent session create --agent my-first-agent -o json | jq -r '.session.id')
    kagent agent invoke --session $SESSION_ID --task "What is 2+2?"
  2. Open the trace in your tracing backend.

    1. Forward the Grafana port, and leave the command running.
      kubectl port-forward -n telemetry svc/kube-prometheus-stack-grafana 3000:80
    2. In your browser, open Grafana at http://localhost:3000, and log in. For the password, see Explore the telemetry in Grafana.
    3. Open Explore, select the Tempo data source, and select the Search query type.
    4. From the Service Name list, select my-first-agent, and run the query. Selecting kagent-controller instead returns the same traces from the controller’s side.
    5. Click a trace to open it.

  3. Review the span tree. The trace starts with the controller’s spans, continues through the Agent Substrate router, and ends with the agent runtime’s spans. The following example shows the spans of one request, with the service that reported each span.

    lf.a2a.v1.A2AService/SendMessage                         kagent-controller
      lf.a2a.v1.A2AService/SendMessage                       kagent-controller
        POST /*                                              agentgateway
          POST                                               agentgateway
            POST /lf.a2a.v1.A2AService/SendMessage           my-first-agent
              a2a.request                                    my-first-agent
                invocation                                   my-first-agent
                  invoke_agent my_first_agent_my_first_harness   my-first-agent
                    generate_content gpt-4.1-mini            my-first-agent
                      HTTP POST                              my-first-agent
    
  4. To narrow a search to one conversation, search by a correlation attribute, such as gen_ai.conversation.id=<context-id>.

Agent Substrate traces

Agent Substrate records traces for its own work, such as scheduling an Actor onto a Worker and restoring it from a snapshot. These traces are separate from the agent request trace. Agent Substrate traces do not share the request trace’s ID, so a request trace does not show how long the Actor took to resume. To investigate a slow start, look up the Agent Substrate traces from the same time window, or read the Actor’s suspend and resume records, which carry the trace ID of each operation.

ServiceReports
atenet-routerRequests that the router receives, and its calls to ateapi to find or resume the Actor for each request.
ateapiActor lifecycle operations, such as create, resume, and suspend, and the scheduling of Actors onto Workers.
ateletWork on a Worker’s node, such as restoring an Actor from a snapshot.
ateom-gvisorWork inside the sandbox that runs the Actor.
atecontrollerReconciliation of Agent Substrate resources, such as WorkerPools.

Agent Substrate exports traces only when its Helm release sets otel.endpoint, and it keeps 1% of its traces by default. To keep more, raise otel.traces.samplingRatio, as the stack guides do. For the steps, see Send Agent Substrate telemetry to the collector.

Traces from a suspended Actor

Agent Substrate checkpointsCheckpointA durable pin on the snapshot that a Session most recently suspended to, and a record of how far its transcript had advanced. Not a new state: tagging copies the snapshot so that Agent Substrate does not collect it, and a second Session can be forked from it.Learn more an Actor as soon as the response body closes, which is sooner than a batching span exporter normally sends its buffer. Spans still in the buffer at that moment freeze inside the snapshotSnapshotThe stored state that an Actor suspends to, held in object storage. Resuming restores the Actor from its most recent snapshot, which is what makes suspending idle agents cheap.Learn more and reach the backend only when the session next resumes, or never at all for a conversation’s last message.

To avoid losing them, the kagent, codex, and claude runtimes flush their span buffer after each A2A handler returns, before the response completes. The flush is unconditional and waits up to three seconds, and no setting changes either. An agent on the byo runtime flushes only if its own image does, so a conversation’s last turn can lose its spans there.

The flush lets a kagent trace arrive promptly rather than on the exporter’s own schedule. To understand what suspension does to an Actor, see Suspend and resume.

Turn tracing off

Turn off the trace exporter, then create a new Session so that the change takes effect.

  1. Disable tracing in the kagent Helm release.

    helm upgrade kagent \
      oci://ghcr.io/kagent-dev/kagent/helm/kagent \
      --version 1.0.0-alpha7 \
      --namespace kagent --reuse-values \
      --set otel.tracing.enabled=false
  2. Create a new Session to pick up the change, because an existing Actor keeps the configuration it started with.

  3. To remove the tracing backend, follow the cleanup steps in the OTel stack or Lightweight OTel stack guide.

Next steps