---
title: "The 2026 evals and observability tools landscape"
description: "Build, buy or combine You can run serious evals with a few Python files, a spreadsheet and CI. As usage grows, dedicated tools save time on tracing…"
url: https://optimizeall.com/learn/llm-evals-and-observability/evals-tools-landscape
updated: 2026-10-05
---

Evaluating and Monitoring LLM Applications · Eval infrastructure: CI regression suites and tools · lesson 10 of 16 · 13 min

# The 2026 evals and observability tools landscape

## Build, buy or combine

You can run serious evals with a few Python files, a spreadsheet and CI. As usage grows, dedicated tools save time on tracing, dataset management, annotation, experiment comparison and dashboards. The landscape moves fast (acquisitions and deprecations happened throughout 2026), so choose on durable criteria and keep your data portable.

## The landscape in 2026 (verify current features and licensing)

| Tool | Type | Notable for |
|---|---|---|
| **promptfoo** | Open-source CLI + web viewer | Declarative test configs, many providers, CI-friendly, red-teaming features; OpenAI announced its acquisition in March 2026 and committed to maintaining open source |
| **Inspect** (UK AI Security Institute) | Open-source Python framework (MIT) | Dataset/solver/scorer abstractions, agent evals with sandboxes, log viewer, many pre-built benchmark evals |
| **Langfuse** | Open-source observability and evals platform, self-hostable or cloud | Tracing, prompt management, datasets, LLM-as-judge evaluators, annotation queues; acquired by ClickHouse in January 2026 and stated to remain open source |
| **LangSmith** (LangChain) | Commercial platform, framework-agnostic | Tracing, datasets, experiments, online evaluators, annotation queues |
| **Braintrust** | Commercial platform | Experiments and eval comparison, datasets, logging, prompt playground |
| **Arize Phoenix** | Open-source observability and evals | OpenTelemetry-native tracing via OpenInference conventions, evals, datasets and experiments |
| **MLflow** | Open-source ML platform | GenAI tracing and evaluation within broader MLOps |
| **Ragas**, **DeepEval** | Open-source metric libraries | RAG and LLM metrics (faithfulness, relevance, and more), pytest-style testing (DeepEval) |
| **OpenAI Evals** | Hosted evals in the OpenAI platform plus the older open-source `openai/evals` repo | OpenAI announced in June 2026 that the hosted Evals platform will become read-only and then shut down (dates in OpenAI's deprecation notice); it suggests migration paths including Promptfoo |

Other capable options exist (for example W&B Weave, Opik, Datadog LLM Observability and cloud providers' own evaluation services). Absence from this table is not a judgment.

## Selection criteria

1. **Data ownership and residency.** Traces contain prompts and user data. Can you self-host? Where is data stored? This matters for clients in the Gulf with residency expectations and for UK/EU GDPR.
2. **OpenTelemetry support.** Tools that ingest standard OTel traces reduce lock-in (next module).
3. **Eval workflow fit.** Code-first (pytest, Inspect, promptfoo) or UI-first (platforms)? Most teams need both: code for CI, UI for annotation and exploration.
4. **Annotation and human review.** Queues, rubrics, inter-annotator agreement.
5. **Framework neutrality.** Works with your stack (raw SDKs, LangGraph, LlamaIndex, OpenAI Agents SDK, Claude Agent SDK, your own code).
6. **Cost model.** Per-seat, per-trace, per-event; self-hosted infrastructure cost.
7. **Vendor stability.** After the 2026 acquisitions and deprecations, check roadmap commitments and export options.

## A pragmatic reference stack

For many teams:

- **Evals as code** in the repo (pytest or promptfoo or Inspect) with datasets in version control.
- **OpenTelemetry instrumentation** in the app, exporting to one observability backend (open-source self-hosted, or commercial).
- **The backend's annotation queues** for human review, with labels exported back into versioned datasets.
- **CI** runs the code-based suite; the platform hosts dashboards and online evaluators.

This keeps your most valuable assets (datasets, graders, instrumentation) portable.

## Hands-on: a minimal Inspect eval for comparison

```python
# support_eval.py: run: inspect eval support_eval.py --model <provider>/<model-id>
from inspect_ai import Task, task
from inspect_ai.dataset import json_dataset, FieldSpec
from inspect_ai.scorer import model_graded_fact
from inspect_ai.solver import generate, system_message

@task
def support_policy():
    return Task(
        dataset=json_dataset("golden_v3.jsonl", FieldSpec(input="input", target="reference_answer", id="id")),
        solver=[system_message(open("prompts/support_system.txt").read()), generate()],
        scorer=model_graded_fact(),
    )
```

Running the same dataset through two tools is a good way to see which workflow your team prefers.

## Worked example

A digital agency in Lahore serving clients in Saudi Arabia chose a self-hosted open-source observability platform because two clients required data to stay in approved cloud regions. Evals ran as pytest in CI; annotation happened in the platform by client-side subject experts; labels were exported weekly into the repo. When they later trialled a commercial platform for a different client, the portable datasets and OTel instrumentation made the pilot a configuration change rather than a rewrite.

## Pitfalls

- **Tool first, process second.** No tool fixes missing error analysis.
- **Lock-in through proprietary SDK calls everywhere.** Wrap them or prefer OTel.
- **Sending sensitive traces to a vendor** without a data processing agreement.
- **Assuming a tool's built-in metrics fit your product.** Calibrate them like any judge.

## How to measure success

Your datasets, graders and instrumentation could move to another tool within a week, and the chosen tool meets your data-residency and review needs.

## Video lecture: The 2026 evals and observability tools landscape

Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.

1. The evals and observability tools landscape
2. Analogy: storing family photos
3. Code-first frameworks
4. Platforms
5. Libraries + a deprecation
6. Seven criteria
7. Reference stack
8. Case: Lahore agency, KSA clients
9. Pitfalls
10. Run a bake-off
11. Example: switching backends in a week
12. Common mistakes
13. Deeper: storage and processing
14. Watch me do it: a one-day bake-off
15. Recap

## Lecture transcript

### The evals and observability tools landscape

In twenty twenty-six alone, the evaluation tool market saw acquisitions, rebrands and a major platform deprecation. If you pick tools by hype, you will be migrating every year. In this lecture you will get a clear map of today's evaluation and observability tools, seven durable criteria for choosing, and a reference stack that keeps your most valuable assets portable.

### Analogy: storing family photos

An analogy for choosing tools. Choosing an eval platform is like choosing where to store your family photos. The app with the nicest interface today might shut down or change its terms next year. People who sleep well keep their originals in a format they control and use apps on top. For evals, your originals are datasets, graders and instrumentation. Keep those portable, and any tool becomes a replaceable app on top.

### Code-first frameworks

Start with code-first frameworks. promptfoo is an open-source command line tool with declarative test files, many providers and red-teaming features; OpenAI announced its acquisition in March twenty twenty-six and committed to maintaining the open-source project. Inspect, from the UK AI Security Institute, is an open-source Python framework with datasets, solvers and scorers, strong agent evaluation with sandboxes, and a great log viewer.

### Platforms

Then platforms for tracing, datasets and human review. Langfuse is open source and self-hostable, with tracing, prompt management, datasets, judges and annotation queues; ClickHouse acquired it in January twenty twenty-six and said it stays open source. LangSmith, from LangChain, is a commercial, framework-agnostic platform. Braintrust is a commercial platform focused on experiments and comparisons. Arize Phoenix is open source and built around OpenTelemetry. MLflow brings GenAI tracing and evaluation into a broader MLOps platform.

### Libraries + a deprecation

Metric libraries like Ragas and DeepEval provide ready-made metrics such as faithfulness and relevance, and DeepEval adds pytest-style testing. And an important change: OpenAI announced in June twenty twenty-six that its hosted Evals platform will become read-only and then shut down, with the dates in its deprecation notice, and it points users toward migration paths including Promptfoo. If you relied on it, plan your move now.

### Seven criteria

Now the seven criteria. One, data ownership and residency, since traces contain user data; can you self-host, and where is data stored? Two, OpenTelemetry support to reduce lock-in. Three, workflow fit: code-first for CI, interface-first for annotation and exploration; most teams need both. Four, annotation and human review features. Five, framework neutrality. Six, the cost model. And seven, vendor stability, including export options.

### Reference stack

Here is a pragmatic reference stack. Evals as code in your repository, using pytest, promptfoo or Inspect, with datasets in version control. OpenTelemetry instrumentation in your application, exporting to one observability backend. That backend's annotation queues for human review, with labels exported back into your versioned datasets. CI runs the code suite, and the platform hosts dashboards and online evaluators. Datasets, graders and instrumentation, your most valuable assets, stay portable.

### Case: Lahore agency, KSA clients

A story from Lahore. A digital agency serving clients in Saudi Arabia chose a self-hosted open-source platform, because two clients required data to stay in approved cloud regions. Evals ran as pytest in CI, client subject experts annotated in the platform, and labels were exported weekly into the repository. Later they trialled a commercial platform for another client. Thanks to portable datasets and OpenTelemetry, the pilot was a configuration change, not a rewrite.

### Pitfalls

Four pitfalls. Choosing a tool before doing error analysis; no tool fixes a missing process. Lock-in through proprietary SDK calls scattered everywhere; wrap them or prefer OpenTelemetry. Sending sensitive traces to a vendor without a data processing agreement. And trusting a tool's built-in metrics without calibrating them against your own labels.

### Run a bake-off

One practical tip before you commit to anything. Run a short bake-off with your real data, not a demo project. Take one feature, thirty test items, and a day of real traces. Instrument once, send the traces to two candidate backends, run your eval suite through two frameworks, and have your domain expert annotate twenty traces in each interface. You will learn more in that day than from any comparison article, including this one.

### Example: switching backends in a week

A simple example of the reference stack in action. Your dataset lives as a JSON lines file in the repository. Your graders are Python functions with unit tests. Your app emits OpenTelemetry traces. On Monday you send traces to one open-source backend. On Friday a client asks for a different, commercial platform. You change an exporter endpoint and credentials, point the platform at the same dataset file, and your graders still run in CI. Nothing important had to be rewritten, because nothing important lived only inside a vendor.

### Common mistakes

Common mistakes when choosing tools. Picking by feature checklists without trying your real data. Ignoring where trace data is stored until a client's security review asks. Scattering one vendor's SDK calls through business logic. And forgetting annotation: the tool your domain experts will actually use to label traces matters as much as the dashboards. Quick question: if your current eval vendor disappeared tomorrow, what would you lose?

### Deeper: storage and processing

One level deeper on data residency. For the Lahore agency's Saudi clients, the question was not only where traces were stored but where they were processed, including any judge model calls made by the platform's online evaluators. Self-hosting the backend solved storage; configuring evaluators to use an approved regional model endpoint solved processing.

### Watch me do it: a one-day bake-off

Watch me do it: a one-day bake-off with the same thirty items. Step one, I keep the dataset as a JSON lines file in the repository. Step two, I run it through our pytest suite, which calls the real application and uses our graders. Pass rate and failing IDs print in the terminal. Step three, I run the same file through the Inspect task from the lesson, which reads it with a field spec mapping input, target and ID, uses our system prompt, and grades with a model-graded fact scorer. I open the log viewer and compare: the two frameworks disagree on three items, and reading them shows the model-graded scorer accepted answers our key-point grader rejected. Step four, traces. Our app already emits OpenTelemetry spans, so I point the exporter at one open-source backend for the morning and a second, commercial backend for the afternoon, only changing endpoint and credentials. Step five, our support lead annotates twenty traces in each interface and times herself: one takes about eleven minutes, the other about eighteen, because of how context is displayed. Step six, I score both backends against the seven criteria, with data residency weighted highest because of two Gulf clients. The decision takes half a page, and none of our datasets, graders or instrumentation had to change.

### Recap

Recap. The landscape changes every quarter, so choose on durable criteria: data control, OpenTelemetry, workflow fit, annotation, neutrality, cost and stability. Keep evals as code, instrument with OpenTelemetry, and export labels back to versioned datasets. Your next step: run the same small dataset through two tools, for example your pytest suite and the Inspect task in the lesson, and write down which workflow your team prefers and why.

## Key takeaways

- Code-first frameworks (promptfoo, Inspect) suit CI; platforms (Langfuse, LangSmith, Braintrust, Arize Phoenix, MLflow) add tracing, datasets and annotation.
- The market changed in 2026: ClickHouse acquired Langfuse, OpenAI announced a Promptfoo acquisition and a hosted Evals deprecation.
- Choose on data ownership and residency, OpenTelemetry support, workflow fit, annotation, neutrality, cost and stability.
- Keep datasets, graders and instrumentation portable.

## Try it

Run the same 30-item dataset through two tools (e.g. your pytest suite and the Inspect task) and write a one-page comparison against the seven criteria.

- [Previous: Regression suites in CI](https://optimizeall.com/learn/llm-evals-and-observability/regression-suites-in-ci)
- [Next: Tracing LLM apps with OpenTelemetry GenAI conventions](https://optimizeall.com/learn/llm-evals-and-observability/tracing-with-opentelemetry-genai)
- [All lessons of Evaluating and Monitoring LLM Applications](https://optimizeall.com/learn/llm-evals-and-observability)
