---
title: "Tracing LLM apps with OpenTelemetry GenAI conventions"
description: "Why tracing is different for LLM apps A single user question to an LLM application can trigger a query rewrite, a vector search, a rerank, two model…"
url: https://optimizeall.com/learn/llm-evals-and-observability/tracing-with-opentelemetry-genai
updated: 2026-10-05
---

Evaluating and Monitoring LLM Applications · Production observability, monitoring and experimentation · lesson 11 of 16 · 15 min

# Tracing LLM apps with OpenTelemetry GenAI conventions

## Why tracing is different for LLM apps

A single user question to an LLM application can trigger a query rewrite, a vector search, a rerank, two model calls, three tool calls and a guardrail check. When the answer is wrong, slow or expensive, you need to see **every step with its inputs, outputs, timing, tokens and cost**. That is distributed tracing, applied to LLM systems.

## OpenTelemetry and the GenAI semantic conventions

**OpenTelemetry (OTel)** is the vendor-neutral standard for traces, metrics and logs. Its **GenAI semantic conventions** define standard names for LLM telemetry so any backend can understand it. Key facts as of September 2026 (verify against the current spec):

- The conventions are still marked **Development** (not yet stable), so attribute names can change; pin the semantic-convention version your instrumentation uses. For example, the provider attribute is now `gen_ai.provider.name` (earlier drafts used a different name).
- In 2026 the GenAI conventions moved out of the main semantic-conventions repository into a dedicated OpenTelemetry GenAI conventions repository.
- **Inference spans** are named `{gen_ai.operation.name} {gen_ai.request.model}` (for example `chat my-model-id`) and require `gen_ai.operation.name` and `gen_ai.provider.name`.
- Common attributes: `gen_ai.request.model`, `gen_ai.response.model`, `gen_ai.usage.input_tokens`, `gen_ai.usage.output_tokens`, `gen_ai.response.finish_reasons`.
- Agent and tool operations include `invoke_agent` and `execute_tool` (with `gen_ai.tool.name`), so multi-step agents appear as nested spans.
- Metrics include `gen_ai.client.token.usage` and `gen_ai.client.operation.duration`.
- **Prompt and response content capture is opt-in**, because it may contain sensitive data.

Instrumentation libraries exist for major SDKs and frameworks (for example OpenTelemetry's own contrib instrumentations, OpenLLMetry by Traceloop and OpenInference by Arize), and many vendors' agent tools emit OTel natively. Backends such as Langfuse, Arize Phoenix, LangSmith, Braintrust, Datadog, Grafana and others can ingest OTLP traces; check each backend's endpoint and auth docs.

## Hands-on: manual GenAI spans in Python

Auto-instrumentation is convenient, but writing one span by hand teaches the model:

```python
# tracing.py: minimal OTel setup + a GenAI chat span (pip install opentelemetry-sdk opentelemetry-exporter-otlp anthropic)
import os
from opentelemetry import trace
from opentelemetry.trace import SpanKind, Status, StatusCode
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from anthropic import Anthropic

provider = TracerProvider(resource=Resource.create({"service.name": "support-bot", "deployment.environment": "prod"}))
provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter()))  # endpoint/headers from OTEL_EXPORTER_OTLP_* env vars
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("support-bot")
client = Anthropic()
MODEL = os.environ["APP_MODEL"]

def chat(messages: list[dict], system: str) -> str:
    with tracer.start_as_current_span(f"chat {MODEL}", kind=SpanKind.CLIENT) as span:
        span.set_attribute("gen_ai.operation.name", "chat")
        span.set_attribute("gen_ai.provider.name", "anthropic")
        span.set_attribute("gen_ai.request.model", MODEL)
        span.set_attribute("gen_ai.request.max_tokens", 800)
        try:
            resp = client.messages.create(model=MODEL, max_tokens=800, system=system, messages=messages)
        except Exception as e:
            span.record_exception(e)
            span.set_status(Status(StatusCode.ERROR))
            span.set_attribute("error.type", type(e).__name__)
            raise
        span.set_attribute("gen_ai.response.model", resp.model)
        span.set_attribute("gen_ai.response.finish_reasons", [resp.stop_reason])
        span.set_attribute("gen_ai.usage.input_tokens", resp.usage.input_tokens)
        span.set_attribute("gen_ai.usage.output_tokens", resp.usage.output_tokens)
        return resp.content[0].text
```

Wrap retrieval and tools in their own spans (a parent span per user request, children for each step), and add your own attributes with a company prefix for business context: `acme.customer_tier`, `acme.country`, `acme.prompt_version`, `acme.retrieval.top_k`. These are what let you slice quality and cost later.

## What to capture, and what not to

| Capture | Why |
|---|---|
| Model, provider, versions of prompt/index/app | Attribute regressions to changes |
| Tokens, latency per step, finish reasons, errors | Cost, performance, truncation detection |
| Retrieved document IDs and scores | RAG diagnosis |
| Tool names, arguments (sanitized), results status | Agent debugging |
| User feedback and session IDs (pseudonymous) | Link quality signals to traces |

**Content capture** (full prompts and responses) is invaluable for debugging and evals, and risky for privacy. Common patterns: capture content in staging always; in production capture for a sampled percentage with PII redaction before export; restrict access; set retention limits consistent with your privacy notices and laws such as UK GDPR or regional data-protection laws in the Gulf and Pakistan.

## Worked example

A food-delivery platform in Karachi saw average latency spike at dinner time. Traces showed the slow step was not the model but a reranker call timing out and retrying. Token attributes revealed a second issue: a prompt change had doubled input tokens by including the full menu. Both were visible in one trace view within minutes; without tracing, the team had been blaming the model provider.

## Pitfalls

- **Logging, not tracing.** Flat logs without parent-child spans make multi-step debugging painful.
- **Unpinned conventions** breaking dashboards when attribute names change.
- **Raw PII in traces** sent to third-party backends without agreements.
- **Missing version attributes**, so you cannot tell which prompt produced which trace.

## How to measure success

Any user complaint can be traced to a full request tree within minutes, with versions, tokens, latency and (appropriately protected) content.

## Video lecture: Tracing LLM apps with OpenTelemetry GenAI conventions

Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.

1. Tracing LLM apps with OpenTelemetry
2. Analogy: parcel tracking
3. Tracing basics
4. Current state (verify)
5. Key attributes
6. Hands-on
7. Add business context
8. Content and privacy
9. Case: Karachi dinner-time latency
10. Example: the 12-second answer (illustrative)
11. Common mistakes
12. Deeper: a content-capture policy
13. Watch me do it: one traced request
14. Recap

## Lecture transcript

### Tracing LLM apps with OpenTelemetry

A single question to your AI assistant might trigger a query rewrite, a vector search, a reranker, two model calls, three tool calls and a safety check. When the answer is wrong, slow or expensive, which step was it? Without tracing, you are guessing. In this lecture you will learn how OpenTelemetry and its generative AI conventions let you see every step, what to capture, and how to protect privacy while doing it.

### Analogy: parcel tracking

An analogy for tracing. Think of a parcel tracking page. You do not just see delivered or not delivered. You see every scan: collected in Lahore, sorted in Karachi, loaded onto a flight, customs in Dubai, out for delivery. When a parcel is late, you know exactly which step slowed it. A trace does that for one request through your AI system: rewrite, retrieval, rerank, model call, tool call, guardrail, each with its time and details.

### Tracing basics

Distributed tracing records a request as a tree of spans. Each span is one step, with a start time, a duration, attributes and a parent. OpenTelemetry is the vendor-neutral standard for traces, metrics and logs, and its generative AI semantic conventions define standard names for LLM telemetry. That means your traces can be understood by many different backends, and you can switch backends without re-instrumenting your code.

### Current state (verify)

A few current facts, which you should verify against the spec because it moves. The generative AI conventions are still marked as development, not stable, so pin the version you use. In twenty twenty-six they moved into their own dedicated repository. Inference spans are named with the operation and model, like chat followed by the model name, and they require an operation name and a provider name attribute.

### Key attributes

Common attributes include the requested and actual model, input and output token counts, and finish reasons, which tell you whether a response was cut off. Agent and tool operations get their own span types, like invoke agent and execute tool, so a multi-step agent appears as a readable tree. And there are standard metrics for token usage and operation duration. One important default: capturing prompt and response content is opt-in, because it may contain sensitive data.

### Hands-on

The lesson walks through writing one span by hand in Python: set up a tracer provider and an exporter that sends OTLP data to your backend, then wrap each model call in a span, set the operation, provider and model, record errors properly, and add token usage and finish reasons from the response. Auto-instrumentation libraries can do this for you, but writing one span manually makes every dashboard you ever see make sense.

### Add business context

Then add business context. Wrap retrieval and tools in child spans under one parent per user request. Add your own attributes with a company prefix: customer tier, country, prompt version, retrieval top k. These are what let you later ask, for example, is our cost per conversation higher for Emirati customers, or did quality drop after prompt version twelve? Without version attributes, you cannot connect a bad trace to the change that caused it.

### Content and privacy

Now privacy. Full prompts and responses are invaluable for debugging and building eval sets, and risky if mishandled. A common pattern: capture content always in staging; in production, capture for a sampled percentage, redact personal data before export, restrict who can view it, and set retention limits that match your privacy notices and laws like UK GDPR or regional data protection laws in the Gulf and Pakistan. And never send raw personal data to a third-party backend without a data processing agreement.

### Case: Karachi dinner-time latency

A worked example. A food-delivery platform in Karachi saw latency spike every evening at dinner time, and the team blamed the model provider. One trace view showed the truth in minutes. The slow step was a reranker call timing out and retrying. And token attributes revealed a second problem: a recent prompt change had doubled input tokens by including the entire menu. Neither issue was the model.

### Example: the 12-second answer (illustrative)

A simple example of reading one trace. A user complains an answer took twelve seconds. You open the trace. The parent span is twelve seconds. Inside it: query rewrite, four hundred milliseconds; retrieval, three hundred; rerank, nine seconds, with two retries; chat span, two seconds with input tokens and finish reason stop. The model was not slow at all. The reranker was timing out. Without spans, you would have blamed the model provider and maybe switched vendors for nothing.

### Common mistakes

Common mistakes with LLM tracing. Logging flat text lines instead of nested spans. Not recording prompt and model versions, so you cannot connect a bad trace to the change that caused it. Capturing full prompts with personal data and sending them to a third party without an agreement. And building dashboards on attribute names that later change because you did not pin the convention version. Try this now: find one slow or wrong answer from last week and see if you can trace it end to end.

### Deeper: a content-capture policy

One level deeper on content capture. In production, the support bot captures full prompts and responses for one percent of traffic, after a redaction step that masks emails, phone numbers and national ID formats before export. Access to those traces is limited to two roles, and they expire after thirty days. Everything else keeps only metadata: tokens, timings, versions and document IDs.

### Watch me do it: one traced request

Watch me do it with the tracing code from the lesson. First, setup. I create a tracer provider with a resource naming the service support bot and the environment prod, add a batch span processor, and an OTLP exporter that reads its endpoint and headers from environment variables, so switching backends never touches code. Second, the span. Inside the chat function I start a span named chat plus the model ID, with kind client. I set three attributes before calling the model: operation name chat, provider name anthropic, and request model. Third, errors. The call sits in a try block; on an exception I record it, set the span status to error, add the error type attribute, and re-raise, so the failure is visible and the caller still sees it. Fourth, the response. I set response model, finish reasons as a list, and input and output token counts from the usage object. Fifth, business context. I wrap the whole request in a parent span and add our own attributes with a company prefix: prompt version v twelve, country AE, topic returns. Then I send one test message and open the backend. The trace shows the parent span, a retrieval child, and the chat child with tokens and finish reason stop. I filter by prompt version, and I can already compare latency between v eleven and v twelve.

### Recap

Recap. Trace, do not just log. Use OpenTelemetry's GenAI conventions, pin their version, and add your own business and version attributes. Capture content carefully with sampling, redaction, access control and retention. Your next step: instrument one model call with a manual span using the lesson code, send it to a backend of your choice, and add prompt version and country attributes.

## Key takeaways

- Trace each request as a tree of spans covering retrieval, model calls, tools and guardrails.
- OTel GenAI conventions are still in Development and moved to a dedicated repo in 2026; pin the version you use.
- Record model, provider, tokens, finish reasons, errors, plus your own version and business attributes.
- Content capture is opt-in: sample, redact, restrict access and set retention in production.

## Try it

Instrument one model call with a manual GenAI span (lesson code), export it to a backend, and add prompt_version and country attributes.

- [Previous: The 2026 evals and observability tools landscape](https://optimizeall.com/learn/llm-evals-and-observability/evals-tools-landscape)
- [Next: Online monitoring, SLOs and drift](https://optimizeall.com/learn/llm-evals-and-observability/online-monitoring-and-slos)
- [All lessons of Evaluating and Monitoring LLM Applications](https://optimizeall.com/learn/llm-evals-and-observability)
