All articles

English

September 22, 20265 min read

Open-Source Agent Observability with OpenTelemetry and Langfuse

Instrument an AI agent end to end with OpenTelemetry semantics and Langfuse: traces, model and tool spans, evaluations, privacy controls, and useful alerts.

Portrait of Tran Kim Dat

Tran Kim Dat

Full-stack Engineer

Abstract distributed trace linking model, tool, retrieval, and analytics signals in a dark observability system

Traditional application monitoring tells you that an endpoint was slow. Agent observability must tell you why the system chose a tool, what that tool returned, how the result changed the next step, and whether the final outcome was useful.

OpenTelemetry provides the vendor-neutral telemetry foundation. Langfuse adds an open-source product layer for LLM traces, prompts, evaluations, datasets, and experiments. Together they let a team investigate behavior without building an observability platform from scratch.

The unit of analysis is the agent run

A request that invokes an agent may contain several model calls, retrievals, tool executions, guardrails, and retries. If each event appears as an unrelated log line, the story is lost. Represent the complete request as a trace and each meaningful operation as a span.

A useful hierarchy looks like this:

  • Agent or workflow span: the full user-visible run.

  • Model span: each generation, including model identifier, token counts, finish reason, and latency.

  • Tool span: validated invocation, result class, duration, and retry count.

  • Retrieval span: query, source class, result count, and relevance signals.

  • Policy or evaluation span: guardrail outcome, score, reason code, and action.

Observability pipeline from a user request through agent trace spans, OpenTelemetry Collector, Langfuse, evaluations, and alerts
Treat one agent run as a trace: model, tool, and retrieval operations become child spans that can be analyzed, evaluated, and alerted on together.

Why OpenTelemetry is the right foundation

OpenTelemetry standardizes traces, metrics, logs, and context propagation. Its generative-AI semantic conventions define attributes for operations such as chat, agent invocation, workflow invocation, tool execution, and retrieval. The conventions are still evolving, so instrument through a small internal adapter instead of scattering experimental attribute names throughout the codebase.

Propagate W3C trace context across HTTP and queue boundaries. When an agent calls an MCP server or starts a background job, the downstream span should remain connected to the original user request. That single decision turns cross-service debugging from guesswork into navigation.

What Langfuse adds

Langfuse ingests traces and presents the LLM-specific details engineers need: generations, prompts, token usage, costs, tool calls, scores, sessions, and users. Its SDKs batch events asynchronously, and its OpenTelemetry-based direction reduces lock-in at the instrumentation layer.

The project is open source under MIT for its core, with enterprise features separated in designated directories. Teams can use the managed service or self-host. Self-hosting provides control over data location and operations, but it also transfers responsibility for upgrades, backups, scaling, authentication, and security.

Design spans for questions you will actually ask

Before adding instrumentation, write the investigation questions:

  • Which tool sequence caused the failure?

  • Did latency come from the model, retrieval, or a downstream API?

  • Which prompt or model version introduced a regression?

  • How often does the agent hit its step or cost budget?

  • Do low user ratings correlate with a particular tool or retrieval source?

  • Which runs required human review, and why?

Then design span names and attributes to answer those questions. Stable low-cardinality dimensions belong in attributes and metrics. Large prompts, responses, and documents require stricter controls.

Do not turn observability into data leakage

Prompts and tool payloads may contain personal data, credentials, source code, or confidential records. Default to metadata-only capture in production. If content is needed for debugging or evaluation, use explicit sampling, redaction, access control, retention limits, and environment separation.

Never record authorization headers, raw API keys, database credentials, or unrestricted tool outputs. Hash or pseudonymize user identifiers where possible. Make content capture a server-side policy that the model cannot override.

Evals complete the picture

Tracing shows the path; evaluation judges the result. Langfuse supports several evaluation patterns: deterministic code checks, human labels, model-based judges, and online or offline evaluation over datasets.

Use the cheapest reliable evaluator for each requirement. Schemas, citations, tool arguments, and forbidden strings are deterministic. Helpfulness or groundedness may need human review or a calibrated model judge. Store the evaluator version and rubric so scores remain interpretable over time.

A strong evaluation set includes normal cases, edge cases, adversarial inputs, dependency failures, stale context, and previously observed incidents. Promote production failures into regression datasets.

Metrics and alerts that matter

A dashboard should connect reliability, quality, and economics:

  • Success and abandonment rate by workflow.

  • End-to-end latency plus model, tool, and retrieval percentiles.

  • Input/output tokens and estimated cost per successful task.

  • Steps, tool calls, and retries per run.

  • Guardrail blocks, approval requests, and stop reasons.

  • Evaluation scores by prompt, model, tool, and release.

Alert on symptoms that require action: a sudden increase in tool errors, loop length, policy blocks, cost per success, or a sustained fall in evaluation score. Avoid alerts on raw token volume without context; a busy product is not necessarily a broken product.

A pragmatic rollout

  1. Instrument one high-value agent workflow as a single trace.

  2. Add child spans for model, tool, retrieval, and guardrail operations.

  3. Send telemetry through an OpenTelemetry Collector so routing and redaction stay outside application code.

  4. Connect Langfuse and verify trace completeness with content capture disabled.

  5. Add a small set of deterministic scores and human labels.

  6. Create one operational dashboard and two actionable alerts.

  7. Review storage cost, access, and retention before increasing sample volume.

Daily.dev discussions about agent harnesses and production readiness consistently place observability beside tools, context, guardrails, and durable execution. That is the right framing: traces are not an afterthought. They are part of the control system.

The outcome

An observable agent is not one that records everything. It is one whose behavior can be reconstructed with the minimum safe evidence. OpenTelemetry supplies portable context and semantics; Langfuse turns that telemetry into a workflow for debugging and improvement. The result is a shorter path from “the agent behaved strangely” to a testable explanation.

Primary sources and further reading

Information checked on September 22, 2026.

Continue reading

More field notes.

View all articles
Application traffic flowing through a policy and telemetry gateway toward several model providers
EnglishSep 22, 2026

AI Gateway vs Direct Provider SDKs in a Production Next.js App

A practical architecture comparison for Next.js teams choosing between direct model-provider SDKs and an AI gateway for routing, failover, policy, and observability.

ai-gateway · ai-sdk · architecture · llm · nextjs

Read article
Three abstract model cores of different sizes connected to code, tools, memory, and infrastructure symbols
EnglishSep 22, 2026

Open-Weight Coding Models in 2026: A Practical Selection Guide

How to evaluate open-weight coding models by task quality, hardware, context, tool use, license, privacy, and operating cost—not benchmark headlines alone.

AI · coding-models · developer-tools · llm · open-source

Read article
Abstract agent loop enclosed by safety barriers, review controls, and a finite resource meter
EnglishSep 22, 2026

Reliable AI Agents: State, Tool Budgets, Guardrails, and Human Review

A production-minded blueprint for AI agents that can recover state, limit tool use, enforce guardrails, escalate risky actions, and stop predictably.

AI Agents · architecture · guardrails · llm · reliability

Read article

Have a product worth building carefully?

Start a conversation