Evaluating and Monitoring LLM ApplicationsProduction observability, monitoring and experimentation · Lesson 11 of 16

Tracing LLM apps with OpenTelemetry GenAI conventions

Article · 15 min · 8 min lecture

Video lecture

Tracing LLM apps with OpenTelemetry GenAI conventions

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

Tracing LLM apps with OpenTelemetry

  • Why tracing
  • GenAI semantic conventions
  • What to capture, what to protect

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why tracing is different for LLM apps

A single user question to an LLM application can trigger a query rewrite, a vector search, a rerank, two model calls, three tool calls and a guardrail check. When the answer is wrong, slow or expensive, you need to see every step with its inputs, outputs, timing, tokens and cost. That is distributed tracing, applied to LLM systems.

OpenTelemetry and the GenAI semantic conventions

OpenTelemetry (OTel) is the vendor-neutral standard for traces, metrics and logs. Its GenAI semantic conventions define standard names for LLM telemetry so any backend can understand it. Key facts as of September 2026 (verify against the current spec):

  • The conventions are still marked Development (not yet stable), so attribute names can change; pin the semantic-convention version your instrumentation uses. For example, the provider attribute is now gen_ai.provider.name (earlier drafts used a different name).
  • In 2026 the GenAI conventions moved out of the main semantic-conventions repository into a dedicated OpenTelemetry GenAI conventions repository.
  • Inference spans are named {gen_ai.operation.name} {gen_ai.request.model} (for example chat my-model-id) and require gen_ai.operation.name and gen_ai.provider.name.
  • Common attributes: gen_ai.request.model, gen_ai.response.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.response.finish_reasons.
  • Agent and tool operations include invoke_agent and execute_tool (with gen_ai.tool.name), so multi-step agents appear as nested spans.
  • Metrics include gen_ai.client.token.usage and gen_ai.client.operation.duration.
  • Prompt and response content capture is opt-in, because it may contain sensitive data.

Instrumentation libraries exist for major SDKs and frameworks (for example OpenTelemetry's own contrib instrumentations, OpenLLMetry by Traceloop and OpenInference by Arize), and many vendors' agent tools emit OTel natively. Backends such as Langfuse, Arize Phoenix, LangSmith, Braintrust, Datadog, Grafana and others can ingest OTLP traces; check each backend's endpoint and auth docs.

Hands-on: manual GenAI spans in Python

Auto-instrumentation is convenient, but writing one span by hand teaches the model:

# tracing.py: minimal OTel setup + a GenAI chat span (pip install opentelemetry-sdk opentelemetry-exporter-otlp anthropic)
import os
from opentelemetry import trace
from opentelemetry.trace import SpanKind, Status, StatusCode
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from anthropic import Anthropic

provider = TracerProvider(resource=Resource.create({"service.name": "support-bot", "deployment.environment": "prod"}))
provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter()))  # endpoint/headers from OTEL_EXPORTER_OTLP_* env vars
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("support-bot")
client = Anthropic()
MODEL = os.environ["APP_MODEL"]

def chat(messages: list[dict], system: str) -> str:
    with tracer.start_as_current_span(f"chat {MODEL}", kind=SpanKind.CLIENT) as span:
        span.set_attribute("gen_ai.operation.name", "chat")
        span.set_attribute("gen_ai.provider.name", "anthropic")
        span.set_attribute("gen_ai.request.model", MODEL)
        span.set_attribute("gen_ai.request.max_tokens", 800)
        try:
            resp = client.messages.create(model=MODEL, max_tokens=800, system=system, messages=messages)
        except Exception as e:
            span.record_exception(e)
            span.set_status(Status(StatusCode.ERROR))
            span.set_attribute("error.type", type(e).__name__)
            raise
        span.set_attribute("gen_ai.response.model", resp.model)
        span.set_attribute("gen_ai.response.finish_reasons", [resp.stop_reason])
        span.set_attribute("gen_ai.usage.input_tokens", resp.usage.input_tokens)
        span.set_attribute("gen_ai.usage.output_tokens", resp.usage.output_tokens)
        return resp.content[0].text

Wrap retrieval and tools in their own spans (a parent span per user request, children for each step), and add your own attributes with a company prefix for business context: acme.customer_tier, acme.country, acme.prompt_version, acme.retrieval.top_k. These are what let you slice quality and cost later.

What to capture, and what not to

CaptureWhy
Model, provider, versions of prompt/index/appAttribute regressions to changes
Tokens, latency per step, finish reasons, errorsCost, performance, truncation detection
Retrieved document IDs and scoresRAG diagnosis
Tool names, arguments (sanitized), results statusAgent debugging
User feedback and session IDs (pseudonymous)Link quality signals to traces

Content capture (full prompts and responses) is invaluable for debugging and evals, and risky for privacy. Common patterns: capture content in staging always; in production capture for a sampled percentage with PII redaction before export; restrict access; set retention limits consistent with your privacy notices and laws such as UK GDPR or regional data-protection laws in the Gulf and Pakistan.

Worked example

A food-delivery platform in Karachi saw average latency spike at dinner time. Traces showed the slow step was not the model but a reranker call timing out and retrying. Token attributes revealed a second issue: a prompt change had doubled input tokens by including the full menu. Both were visible in one trace view within minutes; without tracing, the team had been blaming the model provider.

Pitfalls

  • Logging, not tracing. Flat logs without parent-child spans make multi-step debugging painful.
  • Unpinned conventions breaking dashboards when attribute names change.
  • Raw PII in traces sent to third-party backends without agreements.
  • Missing version attributes, so you cannot tell which prompt produced which trace.

How to measure success

Any user complaint can be traced to a full request tree within minutes, with versions, tokens, latency and (appropriately protected) content.

Key takeaways

  • Trace each request as a tree of spans covering retrieval, model calls, tools and guardrails.
  • OTel GenAI conventions are still in Development and moved to a dedicated repo in 2026; pin the version you use.
  • Record model, provider, tokens, finish reasons, errors, plus your own version and business attributes.
  • Content capture is opt-in: sample, redact, restrict access and set retention in production.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which attributes does an OTel GenAI inference span require?
  2. Why add attributes like prompt_version and country to spans?

Put it into practice

Instrument one model call with a manual GenAI span (lesson code), export it to a backend, and add prompt_version and country attributes.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.