---
title: "Observability: tracing, logging and dashboards for agents"
description: "You cannot fix what you cannot see When an agent misbehaves in production, the first question is always \"what did it see and do?\" Without traces you are…"
url: https://optimizeall.com/learn/ai-agents-engineering/observability-and-tracing
updated: 2026-10-05
---

Building Production AI Agents · Evaluation and observability · lesson 13 of 18 · 15 min

# Observability: tracing, logging and dashboards for agents

## You cannot fix what you cannot see

When an agent misbehaves in production, the first question is always "what did it see and do?" Without traces you are guessing. Observability for agents means capturing, for every run: the inputs, each model call (prompt version, model, parameters, tokens, latency, stop reason), each tool call (arguments, result size, errors, duration), approvals, guardrail decisions and the final outcome, all linked by a single run ID.

## Traces, spans and the agent run

Use the standard tracing model:

- A **trace** = one agent run (or one user request).
- **Spans** = nested operations within it: `agent.run` → `llm.call` (step 0) → `tool.search_crm` → `llm.call` (step 1) → ...
- **Attributes** on spans: model, token counts, tool name, error flags, cost estimate.
- **Events** for notable moments: guardrail triggered, approval requested, compaction happened.

**OpenTelemetry (OTel)** is the vendor-neutral standard. It defines **GenAI semantic conventions** (attributes prefixed `gen_ai.`, such as the requested model and input/output token usage) so traces from different providers and frameworks look alike. The conventions are still evolving, so pin versions and check the current spec. Many agent frameworks emit OTel or their own traces out of the box (the OpenAI Agents SDK traces runs by default to OpenAI's dashboard and supports other processors; LangGraph pairs with LangSmith; the MCP 2026-07-28 spec documents trace-context propagation in `_meta`). Backends range from open-source (Jaeger, Grafana Tempo, Langfuse, Arize Phoenix) to commercial APM tools.

## What to log, and what not to

Log: prompt template IDs and versions, model IDs, parameters, token usage, stop reasons, tool names, argument summaries, errors, latencies, costs, user and tenant IDs (pseudonymized), approval decisions.

Be careful with: full prompts and tool outputs containing personal data. Options: redact PII before logging, store full payloads in a restricted store with short retention, or sample. Align retention with your privacy notices and laws such as GDPR/UK GDPR and Gulf data protection laws; users' messages to an agent are personal data when they identify someone.

## Dashboards that matter

| Panel | Why |
|---|---|
| Runs per hour, success rate, escalation rate | Health at a glance |
| p50/p95 latency per run and per step | Spot slow tools or models |
| Tokens and cost per run, per tenant, per feature | Budget control and pricing |
| Tool error rate by tool | Broken integrations surface fast |
| Steps per run distribution | Loops and regressions show as long tails |
| Guardrail triggers and policy violations | Security signal |
| Stop reasons (max_tokens, refusal, budget exhausted) | Configuration problems |

Alert on: success-rate drops, cost per run spikes, tool error spikes, and any tier-3 action without an approval event (should be impossible; if it happens, page someone).

## Worked example: finding a hidden loop

A UAE real-estate brokerage's lead-qualification agent saw costs double over a week with no traffic change. The steps-per-run histogram showed a new tail at 25+ steps. Traces revealed that a CRM API change started returning an empty list with a 200 status for a renamed field; the agent kept retrying searches with different phrasing. Fixes: the tool now returns an explicit error when the upstream schema is unexpected, and a per-run step budget of 12 triggers escalation. Detection took one dashboard glance; diagnosis took one trace.

## Hands-on: OpenTelemetry spans around your loop

```python
# pip install opentelemetry-sdk opentelemetry-exporter-otlp
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter

provider = TracerProvider()
provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter()))  # endpoint via OTEL_EXPORTER_OTLP_ENDPOINT
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("sales-agent")

def traced_llm_call(client, step: int, **kwargs):
    with tracer.start_as_current_span(f"llm.call step={step}") as span:
        span.set_attribute("gen_ai.operation.name", "chat")
        span.set_attribute("gen_ai.request.model", kwargs["model"])
        resp = client.messages.create(**kwargs)
        span.set_attribute("gen_ai.usage.input_tokens", resp.usage.input_tokens)
        span.set_attribute("gen_ai.usage.output_tokens", resp.usage.output_tokens)
        span.set_attribute("agent.stop_reason", resp.stop_reason)
        return resp

def traced_tool(name: str, fn, args: dict):
    with tracer.start_as_current_span(f"tool.{name}") as span:
        span.set_attribute("agent.tool.name", name)
        try:
            out = fn(**args)
            span.set_attribute("agent.tool.result_chars", len(str(out)))
            return out
        except Exception as exc:
            span.record_exception(exc)
            span.set_attribute("agent.tool.error", True)
            raise

# Wrap the whole run: with tracer.start_as_current_span("agent.run") as run_span: ...
```

Attribute names beginning `gen_ai.` follow the OTel GenAI conventions at the time of writing; check the current spec, as some names have changed between versions. Custom attributes (`agent.*`) are yours to define; document them.

## Replay and debugging

Store enough per step to **replay** a run against a new prompt or model in a sandbox. Replaying failed production traces against a candidate fix is one of the fastest ways to validate improvements, and failed traces should flow into your eval set (lesson 12).

## Pitfalls

- Logging only final answers.
- Logging everything including secrets and full personal data forever.
- No correlation ID across services, so tool-side logs cannot be joined to agent traces.
- Dashboards nobody looks at; tie alerts to on-call.

## Measuring success

Mean time to detect and to diagnose incidents, share of runs with complete traces, and share of production failures converted into eval cases.

## Video lecture: Observability: tracing, logging and dashboards for agents

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. Observability for agents
2. Why it matters
3. Traces, spans, attributes, events
4. OpenTelemetry GenAI conventions
5. What to log
6. Simple example: 'it was slow'
7. Dashboards and alerts
8. Example: the hidden loop
9. Hands-on: OTel spans
10. Pitfalls and metrics
11. Trace retention
12. Deeper: the brokerage loop (illustrative)
13. Watch me do it: OTel spans
14. Try this now
15. Recap

## Lecture transcript

### Observability for agents

When an agent misbehaves in production, the first question is always the same. What did it see, and what did it do? Without traces, you're guessing. In this lesson you'll learn how to instrument agents with traces and spans, what to log and what to protect, and which dashboards and alerts actually matter.

### Why it matters

Why invest in observability before anything goes wrong? Because agent failures are rarely a single crash with a clear error. They're a slow drift: costs creep up, answers get slightly worse, one tool starts timing out. Think of an aircraft's black box and cockpit instruments. Pilots don't wait for an accident to install them. The instruments help them fly every day, and the recorder explains the rare bad day. Traces and dashboards are both for your agents.

### Traces, spans, attributes, events

Use the standard tracing model. A trace is one agent run. Spans are the operations inside it, nested: the run, then each model call, then each tool call, and so on. Spans carry attributes, like the model, token counts, the tool name and whether there was an error. And events mark notable moments: a guardrail fired, an approval was requested, the context was compacted. Everything links through one run id.

### OpenTelemetry GenAI conventions

OpenTelemetry is the vendor neutral standard for this, and it defines generative AI semantic conventions: attribute names starting with gen underscore ai, for things like the requested model and token usage. That means traces from different providers and frameworks look the same in your tools. The conventions are still evolving, so pin versions and check the current spec. Many frameworks emit traces out of the box, and backends range from open source options like Jaeger, Tempo, Langfuse and Phoenix to commercial monitoring tools.

### What to log

What should you log? Prompt template ids and versions, model ids, parameters, token usage, stop reasons, tool names, argument summaries, errors, latencies, costs, pseudonymized user and tenant ids, and approval decisions. Be careful with full prompts and tool outputs, because they often contain personal data. Redact before logging, keep full payloads in a restricted store with short retention, or sample. Align retention with your privacy notices and with laws like GDPR and the Gulf data protection laws.

### Simple example: 'it was slow'

A simple example. A user says the agent was slow this morning. Without tracing, you ask them to try again. With tracing, you search their run id and see a waterfall: the model calls took two seconds each, but one CRM search took forty seconds, three times, because of a timeout and retries. Now you know it's not the model and not the prompt. It's one slow tool. You add a timeout, a cache and an alert. Ten minutes of diagnosis instead of a day of guessing.

### Dashboards and alerts

Now dashboards. Show runs per hour, success rate and escalation rate. Latency at the median and the ninety fifth percentile, per run and per step. Tokens and cost per run, per tenant and per feature. Tool error rate by tool. The distribution of steps per run, where loops show up as long tails. Guardrail triggers and policy violations. And stop reasons, like max tokens or budget exhausted. Then alert on drops in success, cost spikes, tool error spikes, and any high risk action without an approval event. That last one should be impossible, so page someone if it happens.

### Example: the hidden loop

A worked example. A real estate brokerage in the UAE saw its lead qualification agent's costs double in a week, with no extra traffic. The steps per run chart showed a new tail beyond twenty five steps. One trace explained it: a CRM field had been renamed, and the API started returning an empty list with a success status. The agent kept rephrasing its search. The fixes: the tool now returns an explicit error when the data shape is unexpected, and a twelve step budget triggers escalation. Detection took one glance. Diagnosis took one trace.

### Hands-on: OTel spans

The hands on code wraps your model calls and tool calls in OpenTelemetry spans. Each model span records the operation, the model, input and output tokens and the stop reason. Each tool span records the tool name, the result size and any exception. Export over the standard protocol to any backend. Then go one step further: store enough per step to replay a run in a sandbox against a new prompt or model. Replaying failed traces is one of the fastest ways to validate a fix, and those failures should become eval cases.

### Pitfalls and metrics

Avoid four pitfalls: logging only final answers, logging everything including secrets forever, no correlation id across services so tool logs can't be joined to agent traces, and dashboards nobody looks at. Measure your observability with mean time to detect, mean time to diagnose, the share of runs with complete traces, and the share of production failures converted into eval cases.

### Trace retention

How long should you keep traces? Balance debugging needs against cost and privacy. A common pattern: keep full traces, with redacted payloads, for a short window like two weeks, keep aggregated metrics for a year or more, and keep sampled traces of failures and escalations longer for evaluation work. Make sure anything containing personal data follows your retention policy, and that access to raw traces is limited to the people who need it.

### Deeper: the brokerage loop (illustrative)

Let's deepen the UAE brokerage example. Before tracing, the team only had the monthly invoice and a vague sense that the agent felt slower. After adding spans, the dashboard showed the median run had eight steps, but the ninety fifth percentile had jumped to twenty seven after a CRM API change, illustrative figures. The trace of one long run showed the same search repeating with slightly different phrasings, each returning an empty list with a success status. Fixing the tool and adding a step budget pulled the tail back in, and cost per qualified lead returned to normal within a day. Their postmortem's main lesson: an empty success response from a tool is often more dangerous than an error.

### Watch me do it: OTel spans

Watch me do it. Let's walk through the tracing code. I create a tracer provider, attach a batch span processor with the OTLP exporter, whose endpoint comes from an environment variable, and set it globally. Traced llm call opens a span named with the step number, sets the operation name and requested model, makes the real call, then records input tokens, output tokens and the stop reason on the span. Traced tool opens a span per tool with the tool name, runs it, records the result size, and if it throws, records the exception and flags the error before re raising, so my loop's error handling still works. I wrap the whole run in an agent run span. When I run ten tasks against a local tracing backend, each run appears as a waterfall: model calls, tool calls, and one red span where a tool failed.

### Try this now

Try this now. Wrap your agent's model calls and tool calls in spans, using the code in the lesson, and send them to a local tracing tool. Run ten tasks. Then build three panels: success rate, cost per run and steps per run. Finally, pick one failed or slow run and read its trace from top to bottom. Write one sentence explaining what happened. If you can't, add the missing attribute that would have told you.

### Recap

To recap: trace every run with nested spans, use OpenTelemetry conventions, log metadata generously but protect personal data, build dashboards that surface loops, costs and security signals, and replay failures into your evals. Your next step: instrument your agent, export traces to a local backend, and build three panels: success rate, cost per run and steps per run.

## Key takeaways

- Trace every run with nested spans for model calls, tool calls, approvals and guardrails.
- OpenTelemetry's GenAI conventions give vendor-neutral attributes; check the current spec version.
- Log metadata generously but redact or restrict personal data and secrets.
- Dashboards should show success, latency, cost, tool errors, steps distribution and guardrail triggers.
- Replay failed traces and turn them into eval cases.

## Try it

Add spans around the model and tool calls in your agent, export them to a local tracing backend, and build three dashboard panels: success rate, cost per run and steps per run.

- [Previous: Evaluating agents: task success, trajectories and judges](https://optimizeall.com/learn/ai-agents-engineering/evaluating-agents)
- [Next: Cost and latency engineering for agents](https://optimizeall.com/learn/ai-agents-engineering/cost-and-latency-engineering)
- [All lessons of Building Production AI Agents](https://optimizeall.com/learn/ai-agents-engineering)
