Building Production AI AgentsEvaluation and observability · Lesson 13 of 18

Observability: tracing, logging and dashboards for agents

Article · 15 min · 9 min lecture

Video lecture

Observability: tracing, logging and dashboards for agents

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Observability for agents

  • Traces and spans
  • What to log (and not)
  • Dashboards and alerts

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

You cannot fix what you cannot see

When an agent misbehaves in production, the first question is always "what did it see and do?" Without traces you are guessing. Observability for agents means capturing, for every run: the inputs, each model call (prompt version, model, parameters, tokens, latency, stop reason), each tool call (arguments, result size, errors, duration), approvals, guardrail decisions and the final outcome, all linked by a single run ID.

Traces, spans and the agent run

Use the standard tracing model:

  • A trace = one agent run (or one user request).
  • Spans = nested operations within it: agent.run → llm.call (step 0) → tool.search_crm → llm.call (step 1) → ...
  • Attributes on spans: model, token counts, tool name, error flags, cost estimate.
  • Events for notable moments: guardrail triggered, approval requested, compaction happened.

OpenTelemetry (OTel) is the vendor-neutral standard. It defines GenAI semantic conventions (attributes prefixed gen_ai., such as the requested model and input/output token usage) so traces from different providers and frameworks look alike. The conventions are still evolving, so pin versions and check the current spec. Many agent frameworks emit OTel or their own traces out of the box (the OpenAI Agents SDK traces runs by default to OpenAI's dashboard and supports other processors; LangGraph pairs with LangSmith; the MCP 2026-07-28 spec documents trace-context propagation in _meta). Backends range from open-source (Jaeger, Grafana Tempo, Langfuse, Arize Phoenix) to commercial APM tools.

What to log, and what not to

Log: prompt template IDs and versions, model IDs, parameters, token usage, stop reasons, tool names, argument summaries, errors, latencies, costs, user and tenant IDs (pseudonymized), approval decisions.

Be careful with: full prompts and tool outputs containing personal data. Options: redact PII before logging, store full payloads in a restricted store with short retention, or sample. Align retention with your privacy notices and laws such as GDPR/UK GDPR and Gulf data protection laws; users' messages to an agent are personal data when they identify someone.

Dashboards that matter

PanelWhy
Runs per hour, success rate, escalation rateHealth at a glance
p50/p95 latency per run and per stepSpot slow tools or models
Tokens and cost per run, per tenant, per featureBudget control and pricing
Tool error rate by toolBroken integrations surface fast
Steps per run distributionLoops and regressions show as long tails
Guardrail triggers and policy violationsSecurity signal
Stop reasons (max_tokens, refusal, budget exhausted)Configuration problems

Alert on: success-rate drops, cost per run spikes, tool error spikes, and any tier-3 action without an approval event (should be impossible; if it happens, page someone).

Worked example: finding a hidden loop

A UAE real-estate brokerage's lead-qualification agent saw costs double over a week with no traffic change. The steps-per-run histogram showed a new tail at 25+ steps. Traces revealed that a CRM API change started returning an empty list with a 200 status for a renamed field; the agent kept retrying searches with different phrasing. Fixes: the tool now returns an explicit error when the upstream schema is unexpected, and a per-run step budget of 12 triggers escalation. Detection took one dashboard glance; diagnosis took one trace.

Hands-on: OpenTelemetry spans around your loop

# pip install opentelemetry-sdk opentelemetry-exporter-otlp
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter

provider = TracerProvider()
provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter()))  # endpoint via OTEL_EXPORTER_OTLP_ENDPOINT
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("sales-agent")

def traced_llm_call(client, step: int, **kwargs):
    with tracer.start_as_current_span(f"llm.call step={step}") as span:
        span.set_attribute("gen_ai.operation.name", "chat")
        span.set_attribute("gen_ai.request.model", kwargs["model"])
        resp = client.messages.create(**kwargs)
        span.set_attribute("gen_ai.usage.input_tokens", resp.usage.input_tokens)
        span.set_attribute("gen_ai.usage.output_tokens", resp.usage.output_tokens)
        span.set_attribute("agent.stop_reason", resp.stop_reason)
        return resp

def traced_tool(name: str, fn, args: dict):
    with tracer.start_as_current_span(f"tool.{name}") as span:
        span.set_attribute("agent.tool.name", name)
        try:
            out = fn(**args)
            span.set_attribute("agent.tool.result_chars", len(str(out)))
            return out
        except Exception as exc:
            span.record_exception(exc)
            span.set_attribute("agent.tool.error", True)
            raise

# Wrap the whole run: with tracer.start_as_current_span("agent.run") as run_span: ...

Attribute names beginning gen_ai. follow the OTel GenAI conventions at the time of writing; check the current spec, as some names have changed between versions. Custom attributes (agent.*) are yours to define; document them.

Replay and debugging

Store enough per step to replay a run against a new prompt or model in a sandbox. Replaying failed production traces against a candidate fix is one of the fastest ways to validate improvements, and failed traces should flow into your eval set (lesson 12).

Pitfalls

  • Logging only final answers.
  • Logging everything including secrets and full personal data forever.
  • No correlation ID across services, so tool-side logs cannot be joined to agent traces.
  • Dashboards nobody looks at; tie alerts to on-call.

Measuring success

Mean time to detect and to diagnose incidents, share of runs with complete traces, and share of production failures converted into eval cases.

Key takeaways

  • Trace every run with nested spans for model calls, tool calls, approvals and guardrails.
  • OpenTelemetry's GenAI conventions give vendor-neutral attributes; check the current spec version.
  • Log metadata generously but redact or restrict personal data and secrets.
  • Dashboards should show success, latency, cost, tool errors, steps distribution and guardrail triggers.
  • Replay failed traces and turn them into eval cases.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Costs doubled without more traffic. Which chart most quickly points to a looping agent?
  2. What is the main benefit of OpenTelemetry GenAI semantic conventions?
  3. Which logging approach best balances debugging and privacy?

Put it into practice

Add spans around the model and tool calls in your agent, export them to a local tracing backend, and build three dashboard panels: success rate, cost per run and steps per run.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.