{"id":32245,"date":"2026-09-29T13:43:59","date_gmt":"2026-09-29T08:13:59","guid":{"rendered":"https:\/\/opstree.com\/blog\/?p=32245"},"modified":"2026-09-29T13:43:59","modified_gmt":"2026-09-29T08:13:59","slug":"ai-agent-observability-200-ok-wrong-answer","status":"publish","type":"post","link":"https:\/\/opstree.com\/blog\/ai-agent-observability-200-ok-wrong-answer\/","title":{"rendered":"AI Agent Observability: Why 200 OK Doesn&#8217;t Mean Your AI Agent Got It Right"},"content":{"rendered":"<p><span class=\"TextRun SCXW210736507 BCX0\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"auto\"><span class=\"NormalTextRun SCXW210736507 BCX0\">A practical guide to AI agent observability: how tracing, evaluation and structured debugging help teams <\/span><span class=\"NormalTextRun SCXW210736507 BCX0\">monitor<\/span><span class=\"NormalTextRun SCXW210736507 BCX0\"> quality, investigate failures, manage cost, and <\/span><span class=\"NormalTextRun SCXW210736507 BCX0\">operate<\/span><span class=\"NormalTextRun SCXW210736507 BCX0\"> agentic systems more reliably.<\/span><\/span><span class=\"EOP Selected SCXW210736507 BCX0\" data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><b>Quick answer:<\/b> AI agent observability is the practice of capturing an agent\u2019s execution path, model calls, retrieval context, tool inputs and outputs, evaluations, latency, token usage and cost so teams can determine whether the agent completed the user\u2019s task correctly, not merely whether the API returned successfully. <a href=\"https:\/\/opstree.com\/blog\/kubernetes-events-monitoring-using-open-telemetry-and-loki\/\" target=\"_blank\" rel=\"noopener\">OpenTelemetry<\/a> defines traces as the path of a request through an application, with spans representing individual units of work; agent observability applies that tracing idea to LLM calls, tools, retrieval and quality signals. OpenTelemetry traces document the path of a request and model the hierarchy through spans. Langfuse describes tracing as the starting point for ingesting and inspecting LLM and agent execution traces.<\/p>\n<p aria-level=\"2\"><i><span data-contrast=\"none\">Research Basis, Validation Scope and Source Quality<\/span><\/i><span data-ccp-props=\"{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;201341983&quot;:0,&quot;335559738&quot;:360,&quot;335559739&quot;:120,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p>This article should be treated as a research-backed technical guide, not a benchmark report. The conceptual foundation comes from OpenTelemetry\u2019s trace-and-span model, official Langfuse documentation for tracing, masking, datasets, experiments, OpenTelemetry ingestion and self-hosting, plus operational practices from production observability. Any SDK syntax, region availability, licensing detail, hosting dependency or performance number must be checked against the exact Langfuse version and deployment model used at publication time. OpenTelemetry defines traces and spans as the core structure for understanding request execution paths.Langfuse documents OpenTelemetry ingestion through its OTLP endpoint and SDK guidance.Langfuse documents datasets as test collections for evaluating applications.<\/p>\n<h2><span class=\"TextRun SCXW15252387 BCX0\" lang=\"EN\" xml:lang=\"EN\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW15252387 BCX0\" data-ccp-parastyle=\"heading 2\">The Failure Your Dashboard Never Saw<\/span><\/span><\/h2>\n<p><span data-contrast=\"auto\">Consider a support agent in production. Its latency and HTTP error-rate dashboards look healthy. Then a customer reports that the agent claimed a refund had been processed when it had not. The agent called search_orders, received an empty result, retried and eventually produced a confident but unsupported response. The exact numbers and sequence in this scenario are illustrative, the underlying failure pattern is a useful example of a successful request that still delivers an incorrect outcome.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"none\">Your APM recorded that as a successful HTTP request.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"auto\">This is the gap. Traditional application performance monitoring (APM) helps answer whether a service responded and how it performed. Agent observability extends that view, did the agent complete the task correctly, use appropriate tools, rely on relevant evidence and stay within acceptable latency and cost limits? Both operational health and task quality matter.<\/span><\/p>\n<p><span class=\"TextRun SCXW65847886 BCX0\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW65847886 BCX0\">Here&#8217;s<\/span><span class=\"NormalTextRun SCXW65847886 BCX0\"> the same request, seen through both lenses:<\/span><\/span><span class=\"EOP Selected SCXW65847886 BCX0\" data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<div style=\"overflow-x: auto; width: 100%; margin: 25px 0; -webkit-overflow-scrolling: touch;\">\n<table style=\"width: 100%; min-width: 850px; border-collapse: collapse; font-family: Arial,Helvetica,sans-serif; font-size: 14px; line-height: 1.6; color: #374151;\">\n<thead>\n<tr style=\"background: #f5f7fa;\">\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">What you&#8217;re monitoring<\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Traditional APM sees<\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Agent observability sees<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">The request<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">POST \/chat \u2192 200, 4.2s<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Trace with 14 nested steps<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">The work<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">One service call<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">6 LLM calls, 5 tool calls, 3 retrievals<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">The failure<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Nothing &#8211; no exception thrown<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">search_orders returned [] on step 4<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">The recovery<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Invisible<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Agent looped twice, then hallucinated<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">The cost<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Not tracked<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">\u20b918.40 for one conversation<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">The quality<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Not a concept<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Groundedness score: 0.2<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><span class=\"TextRun SCXW247745212 BCX0\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW247745212 BCX0\">Nothing crashed. <\/span><span class=\"NormalTextRun SCXW247745212 BCX0\">That&#8217;s<\/span><span class=\"NormalTextRun SCXW247745212 BCX0\"> exactly the problem.<\/span><\/span><span class=\"EOP Selected SCXW247745212 BCX0\" data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<div style=\"border: 1px solid #d1d5db; padding: 16px; margin: 20px 0; background-color: #f0f4f8;\">\n<p style=\"margin: 0; font-weight: 600; font-size: 16px;\">Also Read: <a href=\"https:\/\/opstree.com\/blog\/enterprise-data-discovery-dpdp-readiness\/\" target=\"_blank\" rel=\"noopener\">Enterprise Data Discovery: Strategy, Tool Selection and DPDP Readiness<\/a><\/p>\n<\/div>\n<h2><span class=\"TextRun SCXW247341242 BCX0\" lang=\"EN\" xml:lang=\"EN\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW247341242 BCX0\" data-ccp-parastyle=\"heading 2\">Why Agent Failures Are Structurally Different<\/span><\/span><span class=\"EOP Selected SCXW247341242 BCX0\" data-ccp-props=\"{&quot;134245418&quot;:false,&quot;134245529&quot;:false,&quot;201341983&quot;:0,&quot;335559738&quot;:300,&quot;335559739&quot;:120,&quot;335559740&quot;:360,&quot;335572079&quot;:6,&quot;335572080&quot;:2,&quot;335572081&quot;:14737632,&quot;469789806&quot;:&quot;single&quot;}\">\u00a0<\/span><\/h2>\n<p><span data-contrast=\"auto\">In conventional services, many failures surface as explicit errors. Agent failures can also be semantic, a tool may return a valid response while the agent misinterprets it, skips a required step or produces an unsupported answer. The final response can sound plausible even when an earlier step was incorrect.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"none\">Three properties make this hard:<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"auto\">Non-determinism. The same input may lead to different outputs or execution paths, depending on model behavior, context, tool state, and system configuration. Capturing the original trace makes investigation more reliable, although replaying a request can still help reproduce some failures.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"none\">Depth, A single user turn fans out into planning, tool selection, retrieval, tool execution, reflection and synthesis. Modern agents add sub-agents and handoffs on top of that. The <\/span><a href=\"https:\/\/opstree.com\/blog\/mcp-agent-is-burning-tokens-before-it-even-starts\/\"><span data-contrast=\"none\">token overhead of loading tools alone<\/span><\/a><span data-contrast=\"none\"> can dominate a request before any reasoning happens.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"none\">Deferred symptoms, The bad tool argument on step 3 doesn&#8217;t surface until the summary on step 11. Without causal structure, you&#8217;re reading a wall of logs trying to work out which line poisoned the well.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"none\">The failure taxonomy that actually shows up in production:<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<div style=\"overflow-x: auto; width: 100%; margin: 25px 0; -webkit-overflow-scrolling: touch;\">\n<table style=\"width: 100%; min-width: 900px; border-collapse: collapse; font-family: Arial,Helvetica,sans-serif; font-size: 14px; line-height: 1.6; color: #374151;\">\n<thead>\n<tr style=\"background: #f5f7fa;\">\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Failure mode<\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">What it looks like<\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Where you catch it<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Wrong tool selected<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Agent calls send_channel_message instead of send_user_message<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Tool call span, input args<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Bad tool arguments<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Correct tool, malformed or hallucinated parameters<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Tool call span, input args<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Silent tool failure<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Tool returns empty\/error, agent proceeds anyway<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Tool output span + downstream generation<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Infinite or near-infinite loop<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Same step repeats 8 times before hitting a cap<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Agent graph view, step counts<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Retrieval miss<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Right question, wrong documents retrieved<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Retriever span, scored for relevance<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Context overflow<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Long session, early instructions fall out of context<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Session view, token counts per turn<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Groundedness failure<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Output not supported by retrieved context<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">LLM-as-a-judge score on the trace<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Cost blowout<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">One user, one session, 40 model calls<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Trace cost aggregation<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><span data-contrast=\"none\">Notice that only two of these throw an exception. The rest need you to look at <\/span><i><span data-contrast=\"none\">content<\/span><\/i><span data-contrast=\"none\">, not status codes.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<h3 aria-level=\"2\"><span data-contrast=\"none\">Traditional APM vs LLM Observability vs Agent Observability<\/span><\/h3>\n<div style=\"overflow-x: auto; width: 100%; margin: 25px 0; -webkit-overflow-scrolling: touch;\">\n<table style=\"width: 100%; min-width: 1100px; border-collapse: collapse; font-family: Arial,Helvetica,sans-serif; font-size: 14px; line-height: 1.6; color: #374151;\">\n<thead>\n<tr style=\"background: #f5f7fa;\">\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Dimension<\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Traditional APM<\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">LLM observability<\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Agent observability<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Primary question<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Did the service respond reliably?<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">What did the model receive and generate?<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Did the agent complete the task correctly across planning, tools, retrieval, and synthesis?<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Core data<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Latency, errors, throughput, infrastructure metrics<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Prompts, completions, tokens, model metadata, cost<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Nested traces, tool calls, retrieval context, evaluations, user feedback, cost attribution<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Failure visibility<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Strong for exceptions and service degradation<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Strong for model input\/output inspection<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Strong for semantic failures, wrong tool use, retrieval misses, loops, and unsupported answers<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Best use<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Operating distributed services<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Improving LLM prompts and outputs<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Operating production AI agents with multiple steps and dependencies<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2><strong>The Three Layers That Make Agents Operable\u00a0<\/strong><\/h2>\n<p><span data-contrast=\"auto\">The rest of this guide follows the same operating loop: first capture the agent\u2019s execution path, then attach quality signals to that path, and finally use the combined evidence to debug, test, and govern future changes.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><em>Agent observability is not one thing. It&#8217;s three layers that feed each other in a loop and skipping any one of them breaks the other two.<\/em><\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-32247 size-large\" src=\"https:\/\/opstree.com\/blog\/wp-content\/uploads\/2026\/09\/agent_observability_diagram_trace_evaluate_debug-1024x398.webp\" alt=\"\" width=\"1024\" height=\"398\" srcset=\"https:\/\/opstree.com\/blog\/wp-content\/uploads\/2026\/09\/agent_observability_diagram_trace_evaluate_debug-1024x398.webp 1024w, https:\/\/opstree.com\/blog\/wp-content\/uploads\/2026\/09\/agent_observability_diagram_trace_evaluate_debug-300x116.webp 300w, https:\/\/opstree.com\/blog\/wp-content\/uploads\/2026\/09\/agent_observability_diagram_trace_evaluate_debug-768x298.webp 768w, https:\/\/opstree.com\/blog\/wp-content\/uploads\/2026\/09\/agent_observability_diagram_trace_evaluate_debug.webp 1388w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/p>\n<p><span data-contrast=\"none\"><strong>Tracing gives you the structure<\/strong> &#8211; a causal record of every step. Without it, you have nothing to evaluate and nothing to debug.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"none\"><strong>Evaluation gives you judgement<\/strong> &#8211; a score attached to that structure, so &#8220;is it good?&#8221; becomes a number you can chart, alert on and gate deploys with.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"none\"><strong>Debugging closes the loop<\/strong> &#8211; you find the broken step, turn that trace into a test case, fix it, and prove the fix with an experiment.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"auto\">The examples use Langfuse as an illustrative observability platform. Product capabilities, SDK APIs, licensing, hosting options, and integrations can change, so verify them against the version you deploy. The underlying practices, trace structure, evaluation, and evidence-led debugging also apply to other <a href=\"https:\/\/buildpiper.io\/glossary\/ai-powered-observability\/\" target=\"_blank\" rel=\"noopener\">observability platforms<\/a> and OpenTelemetry-based architectures, though instrumentation details differ.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<h3><strong><span class=\"TextRun SCXW107849223 BCX0\" lang=\"EN\" xml:lang=\"EN\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW107849223 BCX0\" data-ccp-parastyle=\"heading 2\">Layer 1: Tracing &#8211; Capture What Actually Happened<\/span><\/span><span class=\"EOP Selected SCXW107849223 BCX0\" data-ccp-props=\"{&quot;134245418&quot;:false,&quot;134245529&quot;:false,&quot;201341983&quot;:0,&quot;335559738&quot;:300,&quot;335559739&quot;:120,&quot;335559740&quot;:360,&quot;335572079&quot;:6,&quot;335572080&quot;:2,&quot;335572081&quot;:14737632,&quot;469789806&quot;:&quot;single&quot;}\">\u00a0<\/span><\/strong><\/h3>\n<h4><span class=\"TextRun SCXW110333996 BCX0\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW110333996 BCX0\" data-ccp-parastyle=\"heading 3\">The data model you need to <\/span><span class=\"NormalTextRun SpellingErrorV2Themed SCXW110333996 BCX0\" data-ccp-parastyle=\"heading 3\">internalise<\/span><\/span><span class=\"EOP Selected SCXW110333996 BCX0\" data-ccp-props=\"{&quot;134245418&quot;:false,&quot;134245529&quot;:false,&quot;201341983&quot;:0,&quot;335559738&quot;:180,&quot;335559739&quot;:100,&quot;335559740&quot;:360}\">\u00a0<\/span><\/h4>\n<p><span class=\"TextRun SCXW117923281 BCX0\" lang=\"EN\" xml:lang=\"EN\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW117923281 BCX0\">Everything else depends on getting this right.<\/span><\/span><span class=\"EOP Selected SCXW117923281 BCX0\" data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<div style=\"overflow-x: auto; width: 100%; margin: 25px 0; -webkit-overflow-scrolling: touch;\">\n<table style=\"width: 100%; min-width: 850px; border-collapse: collapse; font-family: Arial,Helvetica,sans-serif; font-size: 14px; line-height: 1.6; color: #374151;\">\n<thead>\n<tr style=\"background: #f5f7fa;\">\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Concept<\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">What it represents<\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Example<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Trace<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">One end-to-end request<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">A user message and the agent&#8217;s full response<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Observation<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">One step inside a trace<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">A single LLM call, tool call or retrieval<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Span<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">A unit of work with duration<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">plan_next_action, execute_tool<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Generation<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">A model call specifically<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Captures prompt, completion, model, tokens, cost<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Session<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Traces grouped into a conversation<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">A 12-turn support chat<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">User<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Traces attributed to a person<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Everything user u_8812 did this week<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Score<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">A quality judgement<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">groundedness: 0.91 on a trace or a single step<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><span class=\"TextRun SCXW192255161 BCX0\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW192255161 BCX0\">Observations nest, That nesting is the whole point, <\/span><span class=\"NormalTextRun SCXW192255161 BCX0\">it&#8217;s<\/span><span class=\"NormalTextRun SCXW192255161 BCX0\"> what lets you see that the bad summary was <\/span><\/span><span class=\"TextRun SCXW192255161 BCX0\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW192255161 BCX0\">caused by<\/span><\/span><span class=\"TextRun SCXW192255161 BCX0\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW192255161 BCX0\"> the empty tool result, rather than just noticing both happened.<\/span><\/span><span class=\"EOP Selected SCXW192255161 BCX0\" data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<h4><span class=\"TextRun SCXW230935814 BCX0\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW230935814 BCX0\" data-ccp-parastyle=\"heading 3\">Why <\/span><span class=\"NormalTextRun SpellingErrorV2Themed SCXW230935814 BCX0\" data-ccp-parastyle=\"heading 3\">OpenTelemetry<\/span><span class=\"NormalTextRun SCXW230935814 BCX0\" data-ccp-parastyle=\"heading 3\"> matters here<\/span><\/span><span class=\"EOP Selected SCXW230935814 BCX0\" data-ccp-props=\"{&quot;134245418&quot;:false,&quot;134245529&quot;:false,&quot;201341983&quot;:0,&quot;335559738&quot;:180,&quot;335559739&quot;:100,&quot;335559740&quot;:360}\">\u00a0<\/span><\/h4>\n<p><span data-contrast=\"auto\">OpenTelemetry matters because it gives teams a common instrumentation path. Langfuse can receive OpenTelemetry traces through its OTLP endpoint, and its documentation recommends using the Langfuse SDK for Python or JavaScript\/TypeScript when available because the SDK handles Langfuse-specific attributes, propagation, media, filtering, and export. This makes it possible to route agent telemetry through familiar observability infrastructure while still using an AI-focused platform for trace inspection, scoring and debugging.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><b><span data-contrast=\"auto\">Architecture flow:<\/span><\/b><span data-contrast=\"auto\"> user request \u2192 agent runtime \u2192 LLM, tool, retrieval, and guardrail spans \u2192 OpenTelemetry SDK or Langfuse SDK \u2192 optional OpenTelemetry collector \u2192 Langfuse for AI-specific trace analysis and evaluation \u2192 existing APM for infrastructure correlation \u2192 dataset and CI regression workflow for future changes. <\/span><span data-contrast=\"auto\">Langfuse documents OTLP ingestion for OpenTelemetry traces.Langfuse datasets support test cases built from inputs and expected outputs.The Langfuse experiment GitHub Action can run experiments in CI and optionally fail a job when regressions are detected.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span class=\"TextRun SCXW8801508 BCX0\" lang=\"EN\" xml:lang=\"EN\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW8801508 BCX0\">This is the difference between observability that your <a href=\"https:\/\/opstree.com\/\" target=\"_blank\" rel=\"noopener\">DevOps team<\/a> adopts and observability that lives on one engineer&#8217;s laptop.<\/span><\/span><span class=\"EOP Selected SCXW8801508 BCX0\" data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<h4><span class=\"TextRun SCXW84618163 BCX0\" lang=\"EN\" xml:lang=\"EN\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW84618163 BCX0\" data-ccp-parastyle=\"heading 3\">Instrumenting: three levels of effort<\/span><\/span><\/h4>\n<div style=\"margin: 25px 0;\">\n<p style=\"margin: 0 0 12px; color: #374151;\"><strong>Level 1 \u2014 drop-in wrapper.<\/strong> Change one import, get traces:<\/p>\n<pre style=\"background: #1e1e1e !important; color: #d4d4d4 !important; padding: 18px; border-radius: 8px; overflow-x: auto; font-family: Consolas,'Courier New',monospace; font-size: 14px; line-height: 1.6; margin: 0 0 28px;\"><code style=\"background: transparent !important; color: inherit !important; padding: 0 !important; margin: 0 !important; border: none !important; box-shadow: none !important;\"># Before\r\n\r\nfrom openai import OpenAI\r\n\r\n# After\r\n\r\nfrom langfuse.openai import openai\r\n\r\ncompletion = openai.chat.completions.create(\r\n    name=\"intent-classification\",\r\n    model=\"gpt-4o\",\r\n    messages=[{\"role\": \"user\", \"content\": user_input}],\r\n    metadata={\"tenant\": \"acme-corp\"},\r\n)<\/code><\/pre>\n<p style=\"margin: 0 0 12px; color: #374151;\"><strong>Level 2 &#8211; framework callback.<\/strong> For LangChain, LangGraph, CrewAI and similar, attach the handler and the framework&#8217;s internal structure becomes your trace structure:<\/p>\n<pre style=\"background: #1e1e1e !important; color: #d4d4d4 !important; padding: 18px; border-radius: 8px; overflow-x: auto; font-family: Consolas,'Courier New',monospace; font-size: 14px; line-height: 1.6; margin: 0 0 28px;\"><code style=\"background: transparent !important; color: inherit !important; padding: 0 !important; margin: 0 !important; border: none !important; box-shadow: none !important;\">from langfuse.langchain import CallbackHandler\r\n\r\nlangfuse_handler = CallbackHandler()\r\n\r\nresponse = agent.invoke(\r\n    {\"messages\": [{\"role\": \"user\", \"content\": \"Where is my refund?\"}]},\r\n    config={\"callbacks\": [langfuse_handler]},\r\n)<\/code><\/pre>\n<p style=\"margin: 0 0 12px; color: #374151;\"><strong>Level 3 &#8211; manual instrumentation.<\/strong> This is where you earn your money. Custom agent loops, business logic, non-LLM steps that still matter:<\/p>\n<pre style=\"background: #1e1e1e !important; color: #d4d4d4 !important; padding: 18px; border-radius: 8px; overflow-x: auto; font-family: Consolas,'Courier New',monospace; font-size: 14px; line-height: 1.6; margin: 0;\"><code style=\"background: transparent !important; color: inherit !important; padding: 0 !important; margin: 0 !important; border: none !important; box-shadow: none !important;\">from langfuse import get_client\r\n\r\nlangfuse = get_client()\r\n\r\nwith langfuse.start_as_current_observation(\r\n    as_type=\"span\", name=\"refund-agent-run\"\r\n) as root:\r\n    root.update(\r\n        input={\"query\": user_query},\r\n        metadata={\"agent_version\": \"2.4.1\", \"tenant\": tenant_id},\r\n    )\r\n\r\n    # Planning step\r\n    with langfuse.start_as_current_observation(\r\n        as_type=\"generation\", name=\"plan\", model=\"gpt-4o\"\r\n    ) as gen:\r\n        plan = call_model(planning_prompt)\r\n        gen.update(output=plan)\r\n\r\n    # Tool execution \u2014 trace the tool, not just the model\r\n    for step in plan.steps:\r\n        with langfuse.start_as_current_observation(\r\n            as_type=\"tool\", name=f\"tool:{step.tool_name}\"\r\n        ) as tool_span:\r\n            tool_span.update(input=step.arguments)\r\n            result = execute_tool(step)\r\n            tool_span.update(\r\n                output=result,\r\n                metadata={\"empty_result\": len(result) == 0},\r\n            )\r\n\r\n    root.update(output=final_answer)\r\n\r\n\r\n# Required in short-lived processes (scripts, serverless, CI)\r\nlangfuse.flush()<\/code><\/pre>\n<\/div>\n<p><span class=\"TextRun SCXW151999939 BCX0\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW151999939 BCX0\">That <\/span><\/span><span class=\"TextRun Highlight SCXW151999939 BCX0\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SpellingErrorV2Themed SCXW151999939 BCX0\">empty_result<\/span><\/span><span class=\"TextRun SCXW151999939 BCX0\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW151999939 BCX0\"> flag is a thirty-second addition that turns an invisible failure into a filterable one. You can now search every trace where a tool came back <\/span><span class=\"NormalTextRun ContextualSpellingAndGrammarErrorV2Themed SCXW151999939 BCX0\">empty<\/span><span class=\"NormalTextRun SCXW151999939 BCX0\"> and the agent answered anyway.<\/span><\/span><\/p>\n<h4><span class=\"TextRun SCXW238164775 BCX0\" lang=\"EN\" xml:lang=\"EN\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW238164775 BCX0\" data-ccp-parastyle=\"heading 3\">What a good trace looks like<\/span><\/span><\/h4>\n<p><span data-contrast=\"none\">Bad traces are worse than no traces, because they create the illusion of visibility. Five rules:<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<ol>\n<li><span data-contrast=\"none\">Name spans semantically, not structurally. <\/span><span data-contrast=\"none\">validate_refund_eligibility<\/span><span data-contrast=\"none\">, not <\/span><span data-contrast=\"none\">step_3<\/span><span data-contrast=\"none\">. Six months from now you&#8217;ll be grateful.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">Capture inputs and outputs at every step, not just at the boundary. The whole value is in the intermediate state.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">Set <\/span><span data-contrast=\"none\">session_id<\/span><span data-contrast=\"none\"> and <\/span><span data-contrast=\"none\">user_id<\/span><span data-contrast=\"none\"> from day one. Retrofitting session grouping onto a live system is miserable.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">Put business context in metadata \u2014 tenant, environment, agent version, feature flag, prompt version. These become your filter dimensions when you&#8217;re triaging.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">Don&#8217;t trace everything. HTTP client spans and database queries from unrelated libraries will drown your agent trace in noise. Instrument deliberately.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:240,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<\/ol>\n<h4><span class=\"TextRun SCXW90485442 BCX0\" lang=\"EN\" xml:lang=\"EN\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW90485442 BCX0\" data-ccp-parastyle=\"heading 3\">The agent graph<\/span><\/span><\/h4>\n<p><span data-contrast=\"auto\">When a trace contains typed observations beyond plain spans, some AI observability platforms can visualize the execution as an agent graph: nodes for steps and edges for control flow. Verify the exact graph capabilities, supported frameworks, and mode names against the platform version you deploy before making product-specific claims in the published blog.Langfuse states that agents can be represented as graphs and that traces can include LLM and non-LLM calls such as retrieval and API calls.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<ul>\n<li aria-setsize=\"-1\" data-leveltext=\"\u25cf\" data-font=\"Roboto\" data-listid=\"1\" data-list-defn-props=\"{&quot;134224900&quot;:false,&quot;201340374&quot;:0,&quot;335551500&quot;:4342338,&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Roboto&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;\u25cf&quot;}\" data-aria-posinset=\"1\" data-aria-level=\"1\"><span data-contrast=\"none\">Aggregated shows the agent&#8217;s overall shape. <\/span><span data-contrast=\"none\">retrieve_docs (3\/3)<\/span><span data-contrast=\"none\"> tells you a step ran three times; loops render as actual cycles. This is the view for &#8220;what does this agent generally do?&#8221;<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<\/ul>\n<ul>\n<li aria-setsize=\"-1\" data-leveltext=\"\u25cf\" data-font=\"Roboto\" data-listid=\"1\" data-list-defn-props=\"{&quot;134224900&quot;:false,&quot;201340374&quot;:0,&quot;335551500&quot;:4342338,&quot;335552541&quot;:1,&quot;335559685&quot;:720,&quot;335559991&quot;:360,&quot;469769226&quot;:&quot;Roboto&quot;,&quot;469769242&quot;:[8226],&quot;469777803&quot;:&quot;left&quot;,&quot;469777804&quot;:&quot;\u25cf&quot;}\" data-aria-posinset=\"2\" data-aria-level=\"1\"><span data-contrast=\"none\">Expanded unrolls every call in execution order. This is the view for &#8220;where exactly did run #4471 go wrong?&#8221;<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:240,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<\/ul>\n<p><span data-contrast=\"none\">For anything with loops, sub-agents or handoffs, this is dramatically faster than scrolling a nested tree. It works with any framework or hand-rolled instrumentation, not just LangGraph.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<h3><span class=\"TextRun SCXW251724660 BCX0\" lang=\"EN\" xml:lang=\"EN\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW251724660 BCX0\" data-ccp-parastyle=\"heading 2\">Layer 2: Evaluation &#8211; Attach Judgement to Structure<\/span><\/span><span class=\"EOP Selected SCXW251724660 BCX0\" data-ccp-props=\"{&quot;134245418&quot;:false,&quot;134245529&quot;:false,&quot;201341983&quot;:0,&quot;335559738&quot;:300,&quot;335559739&quot;:120,&quot;335559740&quot;:360,&quot;335572079&quot;:6,&quot;335572080&quot;:2,&quot;335572081&quot;:14737632,&quot;469789806&quot;:&quot;single&quot;}\">\u00a0<\/span><\/h3>\n<p><span class=\"TextRun SCXW8347020 BCX0\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW8347020 BCX0\">Tracing tells you what happened. It does not tell you whether what happened was <\/span><span class=\"NormalTextRun ContextualSpellingAndGrammarErrorV2Themed SCXW8347020 BCX0\">any good<\/span><span class=\"NormalTextRun SCXW8347020 BCX0\">. <\/span><span class=\"NormalTextRun SCXW8347020 BCX0\">That&#8217;s<\/span><span class=\"NormalTextRun SCXW8347020 BCX0\"> evaluation and it splits cleanly along one axis: are you scoring live traffic, or scoring a change before you ship it?<\/span><\/span><span class=\"EOP Selected SCXW8347020 BCX0\" data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<div style=\"overflow-x: auto; width: 100%; margin: 25px 0; -webkit-overflow-scrolling: touch;\">\n<table style=\"width: 100%; min-width: 850px; border-collapse: collapse; font-family: Arial,Helvetica,sans-serif; font-size: 14px; line-height: 1.6; color: #374151;\">\n<thead>\n<tr style=\"background: #f5f7fa;\">\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\"><\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Online evals<\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Offline evals<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\"><strong>Runs on<\/strong><\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Live production traces<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">A fixed dataset<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\"><strong>Answers<\/strong><\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">&#8220;How are we doing right now?&#8221;<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">&#8220;Is this change better?&#8221;<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\"><strong>Cadence<\/strong><\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Continuous, often sampled<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Per PR, per experiment<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\"><strong>Typical use<\/strong><\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Quality trending, alerting, drift detection<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Regression gates, prompt\/model comparison<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\"><strong>Cost profile<\/strong><\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Scales with traffic \u2014 sample it<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Scales with dataset size \u2014 bounded<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p style=\"margin: 20px 0; font-family: Arial,Helvetica,sans-serif; color: #374151; line-height: 1.7;\"><strong>You need both.<\/strong> Online evals catch the drift you didn&#8217;t predict. Offline evals stop you shipping the regression you did.<\/p>\n<h4><span class=\"TextRun SCXW215898422 BCX0\" lang=\"EN\" xml:lang=\"EN\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW215898422 BCX0\" data-ccp-parastyle=\"heading 3\">Five ways to produce a score<\/span><\/span><\/h4>\n<div style=\"overflow-x: auto; width: 100%; margin: 25px 0; -webkit-overflow-scrolling: touch;\">\n<table style=\"width: 100%; min-width: 900px; border-collapse: collapse; font-family: Arial,Helvetica,sans-serif; font-size: 14px; line-height: 1.6; color: #374151;\">\n<thead>\n<tr style=\"background: #f5f7fa;\">\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Method<\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Best for<\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Trade-off<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Code evaluators<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Deterministic checks &#8211; valid JSON, required fields, PII leakage, length caps<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Cheap, fast, reliable; can&#8217;t judge nuance<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">LLM-as-a-judge<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Groundedness, tone, helpfulness, task completion<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Flexible; costs money, needs calibration<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Human annotation<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Ground truth, ambiguous cases, judge calibration<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Highest quality; doesn&#8217;t scale<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">User feedback<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Real-world signal &#8211; thumbs up\/down, ratings<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Free and honest; sparse and biased<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Custom pipelines<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Domain metrics your business actually cares about<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Full control; you build and maintain it<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><span data-contrast=\"none\">Start with code evaluators. Teams reach for LLM-as-a-judge first because it&#8217;s the interesting one, then discover that 40% of their failures were malformed JSON that a five-line assertion would have caught for free.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><b><span data-contrast=\"none\">Project lesson:<\/span><\/b><span data-contrast=\"none\"> in real implementations, deterministic checks usually deliver the fastest first win. Valid JSON, required fields, empty tool output, policy violations, and schema failures are cheaper and more reliable to detect than subjective quality issues. Add LLM-as-a-judge only after the obvious checks are already automated and calibrated against human review.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<div style=\"margin: 25px 0;\">\n<h3 style=\"margin: 0 0 12px; color: #111827;\">Scoring in practice<\/h3>\n<pre style=\"background: #1e1e1e !important; color: #d4d4d4 !important; padding: 18px; border-radius: 8px; overflow-x: auto; font-family: Consolas,'Courier New',monospace; font-size: 14px; line-height: 1.6; margin: 0;\"><code style=\"background: transparent !important; color: inherit !important; padding: 0 !important; margin: 0 !important; border: none !important; box-shadow: none !important;\">from langfuse import get_client\r\n\r\nlangfuse = get_client()\r\n\r\n\r\n# Attach a score to a trace you already know the ID of\r\nlangfuse.create_score(\r\n    name=\"groundedness\",\r\n    value=0.91,\r\n    trace_id=trace_id,\r\n    data_type=\"NUMERIC\",\r\n    comment=\"All claims supported by retrieved context\",\r\n)\r\n\r\n\r\n# Or score from inside the active context\r\n\r\nwith langfuse.start_as_current_observation(\r\n    as_type=\"span\", name=\"tool-call\"\r\n) as span:\r\n    result = execute_tool(step)\r\n\r\n    # Step-level score \u2014 this is the one people skip\r\n    span.score(\r\n        name=\"tool_returned_data\",\r\n        value=1 if result else 0,\r\n        data_type=\"BOOLEAN\",\r\n    )\r\n\r\n    # And a trace-level score for the run as a whole\r\n    span.score_trace(\r\n        name=\"task_completed\",\r\n        value=1,\r\n        data_type=\"BOOLEAN\",\r\n    )<\/code><\/pre>\n<\/div>\n<p><span data-contrast=\"none\">Score the steps, not just the answer. This is the single highest-leverage habit in agent evaluation. A trace-level score of 0.4 tells you the run was bad. Step-level scores tell you <\/span><i><span data-contrast=\"none\">retrieval was fine, tool execution was fine, synthesis was bad<\/span><\/i><span data-contrast=\"none\"> ,\u00a0 which is the difference between a week of guessing and an afternoon of fixing.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"none\">For user feedback, capture it as a score against the same trace ID and you get a direct join between &#8220;the user was unhappy&#8221; and &#8220;here is the exact reasoning chain that made them unhappy.&#8221;<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<h3><span class=\"TextRun SCXW238040671 BCX0\" lang=\"EN\" xml:lang=\"EN\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW238040671 BCX0\" data-ccp-parastyle=\"heading 3\">Datasets and experiments: the regression gate<\/span><\/span><span class=\"EOP Selected SCXW238040671 BCX0\" data-ccp-props=\"{&quot;134245418&quot;:false,&quot;134245529&quot;:false,&quot;201341983&quot;:0,&quot;335559738&quot;:180,&quot;335559739&quot;:100,&quot;335559740&quot;:360}\">\u00a0<\/span><\/h3>\n<p><span data-contrast=\"none\">Every good production failure should become a permanent test case. The workflow:<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<ol>\n<li><span data-contrast=\"none\">A trace fails in production.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">You add it to a dataset with the expected output.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">Every prompt change, model swap or code change runs against that dataset.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">Scores are compared to the previous run, side by side.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">If a score drops below threshold, the build fails.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:240,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<\/ol>\n<p><span data-contrast=\"none\">That last step is what makes this engineering rather than vibes. Langfuse ships a GitHub Action for experiments on <\/span><span data-contrast=\"none\">pull_request<\/span><span data-contrast=\"none\">; your experiment script raises a regression error when a score violates its threshold and the job fails. Same idea as a failing unit test, applied to a non-deterministic system.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"none\">Build your dataset from three sources: real production failures (highest value), edge cases you can reason about in advance, and a boring happy-path set so you notice when a &#8220;safe&#8221; change breaks the basics.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<div style=\"border: 1px solid #d1d5db; padding: 16px; margin: 20px 0; background-color: #f0f4f8;\">\n<p style=\"margin: 0; font-weight: 600; font-size: 16px;\">Also Read: <a href=\"https:\/\/opstree.com\/blog\/unified-business-intelligence-platform\/\" target=\"_blank\" rel=\"noopener\">From Fragmented Data to Faster Decisions: Building a Unified Business Intelligence Platform<\/a><span style=\"background-color: #ffffff;\">\u00a0<\/span><\/p>\n<\/div>\n<h3><span class=\"TextRun SCXW153959670 BCX0\" lang=\"EN\" xml:lang=\"EN\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW153959670 BCX0\" data-ccp-parastyle=\"heading 2\">Layer 3: Debugging &#8211; Close the Loop<\/span><\/span><\/h3>\n<p><span data-contrast=\"none\">With traces and scores in place, debugging stops being archaeology and becomes a routine.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"none\">The loop:<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<ol>\n<li><span data-contrast=\"none\">Start with a signal, not a hunch. A dropping score, an alert, a cluster of thumbs-down, a cost spike.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">Filter to the population. Traces where <\/span><span data-contrast=\"none\">groundedness &lt; 0.5<\/span><span data-contrast=\"none\"> and <\/span><span data-contrast=\"none\">tenant = acme-corp<\/span><span data-contrast=\"none\"> and <\/span><span data-contrast=\"none\">agent_version = 2.4.1<\/span><span data-contrast=\"none\">. You&#8217;re looking for a pattern, not an anecdote.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">Open the graph view. Find the shape of the failure \u2014 where does the path diverge from healthy runs? Loops and repeated steps announce themselves immediately.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">Drill into the first bad step. Not the bad output \u2014 the first step where the input was fine and the output wasn&#8217;t. That&#8217;s your actual bug.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">Reproduce in the playground. Take the exact prompt and context from that span, change one variable, see what happens.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">Turn it into a dataset item with the correct expected output.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">Fix, then run the experiment. Prove the fix works on that case and doesn&#8217;t break the other forty.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:240,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<\/ol>\n<p><span data-contrast=\"none\">Common symptoms and where to look:<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<div style=\"overflow-x: auto; width: 100%; margin: 25px 0; -webkit-overflow-scrolling: touch;\">\n<table style=\"width: 100%; min-width: 850px; border-collapse: collapse; font-family: Arial,Helvetica,sans-serif; font-size: 14px; line-height: 1.6; color: #374151;\">\n<thead>\n<tr style=\"background: #f5f7fa;\">\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Symptom<\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">First place to look<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Confident but wrong answers<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Retriever span output, then groundedness score on the generation<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Latency spike, no error<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Step count per trace \u2014 the agent is probably looping<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Cost spike<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Token counts per generation; check whether tool definitions are bloating every call<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Works in dev, fails in prod<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Metadata diff \u2014 prompt version, model version, tool availability<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Degrades over a long conversation<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Session view, token count per turn, context window pressure<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Intermittent tool failures<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Tool spans filtered by empty\/error output<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><span class=\"TextRun SCXW74518833 BCX0\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW74518833 BCX0\">The move that pays for the entire setup: compare a failing trace against a passing trace for the same task. Two tabs, same structure, and the divergence point <\/span><span class=\"NormalTextRun ContextualSpellingAndGrammarErrorV2Themed SCXW74518833 BCX0\">is<\/span><span class=\"NormalTextRun SCXW74518833 BCX0\"> usually obvious within a minute.<\/span><\/span><span class=\"EOP Selected SCXW74518833 BCX0\" data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<h2><span class=\"TextRun SCXW59964712 BCX0\" lang=\"EN\" xml:lang=\"EN\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW59964712 BCX0\" data-ccp-parastyle=\"heading 2\">What to Consider Before You Roll This Out<\/span><\/span><span class=\"EOP Selected SCXW59964712 BCX0\" data-ccp-props=\"{&quot;134245418&quot;:false,&quot;134245529&quot;:false,&quot;201341983&quot;:0,&quot;335559738&quot;:300,&quot;335559739&quot;:120,&quot;335559740&quot;:360,&quot;335572079&quot;:6,&quot;335572080&quot;:2,&quot;335572081&quot;:14737632,&quot;469789806&quot;:&quot;single&quot;}\">\u00a0<\/span><\/h2>\n<p><span class=\"NormalTextRun SCXW17794514 BCX0\">Getting a trace into a dashboard is the easy part. <\/span><span class=\"NormalTextRun SCXW17794514 BCX0\">Here&#8217;s<\/span><span class=\"NormalTextRun SCXW17794514 BCX0\"> what separates a demo from something your team relies on at 2am.<\/span><\/p>\n<h3><span class=\"TextRun SCXW76435814 BCX0\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW76435814 BCX0\" data-ccp-parastyle=\"heading 3\">1. Decide what <\/span><span class=\"NormalTextRun SCXW76435814 BCX0\" data-ccp-parastyle=\"heading 3\">you&#8217;re<\/span><span class=\"NormalTextRun SCXW76435814 BCX0\" data-ccp-parastyle=\"heading 3\"> allowed to capture<\/span><\/span><span class=\"EOP Selected SCXW76435814 BCX0\" data-ccp-props=\"{&quot;134245418&quot;:false,&quot;134245529&quot;:false,&quot;201341983&quot;:0,&quot;335559738&quot;:180,&quot;335559739&quot;:100,&quot;335559740&quot;:360}\">\u00a0<\/span><\/h3>\n<p><span data-contrast=\"none\">Agent traces contain prompts and completions, which means they contain whatever your users typed. That&#8217;s PII, and possibly regulated data. Before you instrument anything:<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<ul>\n<li><span data-contrast=\"none\">Use SDK-level masking to redact sensitive fields before they leave your process.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"auto\">Choose a data region deliberately if using Langfuse Cloud. Current documentation lists US, EU, Japan, and HIPAA regions, with accounts and data separated between regions; verify availability and compliance terms before go-live.Langfuse documentation lists Cloud regions and explains that data and accounts are separated between regions.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559740&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">Set retention policies that match your compliance posture, not the default.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">Get this signed off before go-live, not during the audit.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:240,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<\/ul>\n<p><span data-contrast=\"none\">This is the same discipline that applies to <\/span><a href=\"https:\/\/opstree.com\/blog\/secure-enterprise-mcp-server-for-generative-ai\/\"><span data-contrast=\"none\">building a secure enterprise MCP server<\/span><\/a><span data-contrast=\"none\"> \u2014 the observability layer sees everything the agent sees, so it inherits the agent&#8217;s entire threat model.<\/span><\/p>\n<h3><span class=\"TextRun SCXW165150209 BCX0\" lang=\"EN\" xml:lang=\"EN\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW165150209 BCX0\" data-ccp-parastyle=\"heading 3\">2. Sample, and sample intelligently<\/span><\/span><span class=\"EOP Selected SCXW165150209 BCX0\" data-ccp-props=\"{&quot;134245418&quot;:false,&quot;134245529&quot;:false,&quot;201341983&quot;:0,&quot;335559738&quot;:180,&quot;335559739&quot;:100,&quot;335559740&quot;:360}\">\u00a0<\/span><\/h3>\n<p><span data-contrast=\"none\">Tracing every request at full fidelity is affordable at 1,000 requests\/day and painful at 10 million. But naive random sampling is the wrong answer, because failures are rare and random sampling throws most of them away.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"none\">A pattern that works:<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<div style=\"overflow-x: auto; width: 100%; margin: 25px 0; -webkit-overflow-scrolling: touch;\">\n<table style=\"width: 100%; min-width: 750px; border-collapse: collapse; font-family: Arial,Helvetica,sans-serif; font-size: 14px; line-height: 1.6; color: #374151;\">\n<thead>\n<tr style=\"background: #f5f7fa;\">\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Traffic class<\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Sampling rate<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Errors and exceptions<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">100%<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Traces with a failing score<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">100%<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Traces above a cost or latency threshold<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">100%<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">New agent version, first 24h<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">100%<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Everything else<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">1\u201310%<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><span class=\"TextRun SCXW1698330 BCX0\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"auto\"><span class=\"NormalTextRun SCXW1698330 BCX0\">LLM-as-a-judge evaluation can add meaningful cost and latency, depending on model, sampling rate and trace volume. Measure the cost of evaluation and trace storage in your own workload; <\/span><span class=\"NormalTextRun SCXW1698330 BCX0\">retain<\/span><span class=\"NormalTextRun SCXW1698330 BCX0\"> enough evidence to investigate failures.<\/span><\/span><span class=\"EOP Selected SCXW1698330 BCX0\" data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<h3><span class=\"TextRun SCXW192905703 BCX0\" lang=\"EN\" xml:lang=\"EN\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW192905703 BCX0\" data-ccp-parastyle=\"heading 3\">3. Never block the request path<\/span><\/span><\/h3>\n<p><span data-contrast=\"auto\">Observability should not add meaningful latency to the user-facing request path. Use asynchronous batching where available, measure overhead in your own workload, and verify that telemetry failures do not block responses. For short-lived processes such as serverless functions, CLI tools, and CI jobs, explicitly flush before exit so traces are not lost.Langfuse data-region documentation notes that tracing ingestion is sent asynchronously in batches, making ingestion latency less directly relevant to application performance.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<ul>\n<li><span data-contrast=\"none\">In short-lived processes &#8211; serverless functions, CLI tools, CI jobs, call <\/span><span data-contrast=\"none\">flush()<\/span><span data-contrast=\"none\"> before exit or you&#8217;ll silently lose traces.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">If the observability backend is down, your agent must keep serving. Verify this explicitly; don&#8217;t assume it.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:240,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<\/ul>\n<h3><strong><span class=\"TextRun SCXW69392990 BCX0\" lang=\"EN\" xml:lang=\"EN\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW69392990 BCX0\" data-ccp-parastyle=\"heading 3\">4. Treat span naming as a public API<\/span><\/span><span class=\"EOP Selected SCXW69392990 BCX0\" data-ccp-props=\"{&quot;134245418&quot;:false,&quot;134245529&quot;:false,&quot;201341983&quot;:0,&quot;335559738&quot;:180,&quot;335559739&quot;:100,&quot;335559740&quot;:360}\">\u00a0<\/span><\/strong><\/h3>\n<p><span data-contrast=\"none\">Your span names become your filter dimensions, your dashboard groupings and your alert conditions. Rename them casually and you break six months of historical comparison.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<ul>\n<li><span data-contrast=\"none\">\u2705 <\/span><span data-contrast=\"none\">retrieval.search_knowledge_base<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">\u2705 <\/span><span data-contrast=\"none\">tool.jira_create_issue<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">\u2705 <\/span><span data-contrast=\"none\">llm.synthesize_answer<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">\u274c <\/span><span data-contrast=\"none\">step_2<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">\u274c <\/span><span data-contrast=\"none\">call_model<\/span><span data-contrast=\"none\"> (which model? for what?)<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:240,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<\/ul>\n<p><span data-contrast=\"none\">Namespace them. Version them if you must change them. Same rules as <\/span><a href=\"https:\/\/opstree.com\/blog\/mcp-agent-is-burning-tokens-before-it-even-starts\/\"><span data-contrast=\"none\">tool naming in MCP<\/span><\/a><span data-contrast=\"none\"> , it&#8217;s a contract, not a label.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<h3><span class=\"EOP Selected SCXW192905703 BCX0\" data-ccp-props=\"{&quot;134245418&quot;:false,&quot;134245529&quot;:false,&quot;201341983&quot;:0,&quot;335559738&quot;:180,&quot;335559739&quot;:100,&quot;335559740&quot;:360}\"> <span class=\"TextRun SCXW140561347 BCX0\" lang=\"EN\" xml:lang=\"EN\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW140561347 BCX0\" data-ccp-parastyle=\"heading 3\">5. Link prompt versions to traces<\/span><\/span><span class=\"EOP Selected SCXW140561347 BCX0\" data-ccp-props=\"{&quot;134245418&quot;:false,&quot;134245529&quot;:false,&quot;201341983&quot;:0,&quot;335559738&quot;:180,&quot;335559739&quot;:100,&quot;335559740&quot;:360}\">\u00a0<\/span><\/span><\/h3>\n<p><span data-contrast=\"none\">If you&#8217;re managing prompts centrally and you should be record which prompt version produced each generation. Without it, &#8220;quality dropped on Tuesday&#8221; is unanswerable. With it, it&#8217;s a two-click diff between version 7 and version 8.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<h3><span data-contrast=\"none\">6. Evaluate your evaluator<\/span><\/h3>\n<p><span data-contrast=\"none\">LLM-as-a-judge is a model call, which means it has all the same failure modes as the thing it&#8217;s judging. Judges drift, judges are biased toward verbose answers, judges score their own model family higher.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"none\">Calibrate: have a human annotate 50\u2013100 traces, compare against your judge&#8217;s scores, and measure the agreement. If agreement is poor, fix the judge prompt before you trust a single dashboard built on it. Re-check after any judge model upgrade.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<h3><span data-contrast=\"none\">7. Plan the self-hosting reality<\/span><\/h3>\n<p><span data-contrast=\"none\">Langfuse is MIT-licensed and self-hostable on every tier, which is often the deciding factor for enterprise and regulated workloads. Two things to know going in:<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<ul>\n<li><span data-contrast=\"auto\">Self-hosting requires planning for ClickHouse as the main OLAP storage layer for traces, observations, and scores, alongside the other platform dependencies. Langfuse documentation describes ClickHouse as the primary analytical store for these entities and lists supported deployment options and version requirements.Langfuse documents ClickHouse as the main OLAP storage solution for trace, observation, and score entities.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559740&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"auto\">Review the platform&#8217;s current ownership, roadmap, hosting model, and storage dependencies during procurement. These details can change and should be verified against current vendor documentation.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:240,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<\/ul>\n<p><span data-contrast=\"auto\">Benchmark the version you intend to deploy using representative trace volume, retention, query patterns, and concurrency rather than relying on older performance reports.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><b><span data-contrast=\"none\">Enterprise implementation checklist:<\/span><\/b><span data-contrast=\"none\"> validate storage dependencies, backup and restore, retention cost, masking strategy, access controls, tenant isolation, audit requirements, data residency, upgrade process, and the team that owns the platform after launch. Treat agent traces as sensitive operational data because they may contain prompts, completions, retrieved documents, tool arguments, and user-provided information.Langfuse self-hosting documentation describes client-side and server-side masking approaches for sensitive data.Langfuse masking documentation explains how masking functions can redact sensitive tracing data before export.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<h3><span data-contrast=\"none\">8. Make it part of the platform, not a side project<\/span><\/h3>\n<p><span data-contrast=\"none\">The failure mode for observability projects is that one enthusiastic engineer instruments one service, and nobody else adopts it. Avoid it by:<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<ul>\n<li><span data-contrast=\"none\">Putting instrumentation in your shared agent scaffolding, so new agents are traced by default.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">Standardising metadata keys across teams (<\/span><span data-contrast=\"none\">tenant<\/span><span data-contrast=\"none\">, <\/span><span data-contrast=\"none\">env<\/span><span data-contrast=\"none\">, <\/span><span data-contrast=\"none\">agent_version<\/span><span data-contrast=\"none\">) so dashboards work across services.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">Wiring the CI regression gate on day one , that&#8217;s what makes evals load-bearing rather than decorative.<\/span><span data-ccp-props=\"{&quot;134233118&quot;:false,&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:0,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<li><span data-contrast=\"none\">Routing agent spans through your existing OTel collector so this lives alongside your other telemetry instead of beside it.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559685&quot;:720,&quot;335559739&quot;:240,&quot;335559740&quot;:360,&quot;335559991&quot;:360}\">\u00a0<\/span><\/li>\n<\/ul>\n<h2><span class=\"TextRun SCXW99226998 BCX0\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW99226998 BCX0\" data-ccp-parastyle=\"heading 2\">Decision Guide: How Much Observability Do You Actually Need?<\/span><\/span><\/h2>\n<div style=\"overflow-x: auto; width: 100%; margin: 25px 0; -webkit-overflow-scrolling: touch;\">\n<table style=\"width: 100%; min-width: 1000px; border-collapse: collapse; font-family: Arial,Helvetica,sans-serif; font-size: 14px; line-height: 1.6; color: #374151;\">\n<thead>\n<tr style=\"background: #f5f7fa;\">\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Where you are<\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Tracing<\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Online evals<\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Datasets + CI gate<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Prototype, single developer<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Basic \u2014 SDK wrapper<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Skip<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Skip<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Internal tool, low stakes<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Full manual instrumentation<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Sampled, 1\u20132 metrics<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Small dataset, manual runs<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Customer-facing, single agent<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Full + sessions + users<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Sampled + user feedback<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Required<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Multi-agent \/ high stakes<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Full + graph view + step scores<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Comprehensive<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Required + blocking CI<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Regulated (finance, health)<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Full + masking + self-host<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Comprehensive + audit trail<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Required + human annotation<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2 aria-level=\"2\"><span data-contrast=\"none\">Start Here: The First Week<\/span><span data-ccp-props=\"{&quot;134245418&quot;:false,&quot;134245529&quot;:false,&quot;201341983&quot;:0,&quot;335559738&quot;:300,&quot;335559739&quot;:120,&quot;335559740&quot;:360,&quot;335572079&quot;:6,&quot;335572080&quot;:2,&quot;335572081&quot;:14737632,&quot;469789806&quot;:&quot;single&quot;}\">\u00a0<\/span><\/h2>\n<p><span data-contrast=\"none\">If you&#8217;re instrumenting an existing agent, this order gets you value fastest:<\/span><\/p>\n<div style=\"overflow-x: auto; width: 100%; margin: 25px 0; -webkit-overflow-scrolling: touch;\">\n<table style=\"width: 100%; min-width: 750px; border-collapse: collapse; font-family: Arial,Helvetica,sans-serif; font-size: 14px; line-height: 1.6; color: #374151;\">\n<thead>\n<tr style=\"background: #f5f7fa;\">\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Day<\/th>\n<th style=\"border: 1px solid #ddd; padding: 12px; text-align: left;\">Do this<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">1<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Add the SDK wrapper or framework callback. Get <em>any<\/em> trace flowing.<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">2<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Add session_id, user_id and metadata (env, version, tenant).<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">3<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Manually instrument tool calls \u2014 inputs, outputs, and an empty-result flag.<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">4<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Add two code evaluators: output validity, and a &#8220;tool returned data&#8221; check.<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">5<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Pull 20 real traces into a dataset. Include your worst known failures.<\/td>\n<\/tr>\n<tr style=\"background: #fafafa;\">\n<td style=\"border: 1px solid #ddd; padding: 12px;\">6<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Wire up one LLM-as-a-judge metric. Calibrate it against 30 human-labelled traces.<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">7<\/td>\n<td style=\"border: 1px solid #ddd; padding: 12px;\">Put the experiment run in CI with a score threshold.<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><span data-contrast=\"auto\">A focused first week can establish the foundations for better visibility. Detection time will depend on instrumentation coverage, evaluation cadence, alerting, and operational ownership.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p aria-level=\"2\"><span data-contrast=\"none\">Original Value: Practical Lessons That Make This More Than a Tool Overview<\/span><span data-ccp-props=\"{&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;201341983&quot;:0,&quot;335559738&quot;:360,&quot;335559739&quot;:120,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"none\">The most useful agent observability programs do not start with dashboards; they start with failure review. Pick five recent bad answers, trace each one to the first incorrect step, and ask what signal would have caught it earlier. That exercise usually produces a better instrumentation plan than copying a generic observability checklist.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"none\">Common mistakes include tracing only the final response, hiding tool inputs for convenience, failing to version prompts, sampling away rare failures, using LLM-as-a-judge without calibration, and treating cost as a monthly bill instead of a trace-level debugging signal. The highest-value improvement is usually not a new model; it is better evidence about where the current agent failed.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<h2 aria-level=\"2\"><span data-contrast=\"none\">One Last Thing<\/span><span data-ccp-props=\"{&quot;134245418&quot;:false,&quot;134245529&quot;:false,&quot;201341983&quot;:0,&quot;335559738&quot;:300,&quot;335559739&quot;:120,&quot;335559740&quot;:360,&quot;335572079&quot;:6,&quot;335572080&quot;:2,&quot;335572081&quot;:14737632,&quot;469789806&quot;:&quot;single&quot;}\">\u00a0<\/span><\/h2>\n<p><span data-contrast=\"none\">The instinct when an agent misbehaves is to reach for the prompt. Add a line. Tell it to be more careful. Ship it and hope.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"none\">That instinct is the problem. It treats a system with a dozen moving parts as a single text box, and it produces the specific kind of codebase where nobody can explain why the prompt says what it says, and nobody dares change it.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"none\">Observability replaces that with something ordinary and unglamorous: look at what happened, measure whether it was good, find the step that broke, fix that step, prove it with a test. It&#8217;s the same discipline we already apply to distributed systems. Agents don&#8217;t get an exemption just because the failure mode is a paragraph of fluent English instead of a stack trace.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<p><span data-contrast=\"auto\">Teams operating reliable agents need more than carefully written prompts: they need evidence of what happened, a way to measure quality, and a repeatable process for locating and validating fixes.<\/span><span data-ccp-props=\"{&quot;201341983&quot;:0,&quot;335559738&quot;:120,&quot;335559739&quot;:180,&quot;335559740&quot;:360}\">\u00a0<\/span><\/p>\n<h2><span data-ccp-props=\"{}\">Related Searches<\/span><\/h2>\n<ul>\n<li><a href=\"https:\/\/opstree.com\/blog\/leading-telecom-enterprise-transformed-enterprise-analytics-fractal-gpt\/\" target=\"_blank\" rel=\"noopener\">How Leading Telecom Enterprise Transformed Enterprise Analytics with Fractal GPT<\/a><\/li>\n<li><a href=\"https:\/\/opstree.com\/blog\/real-time-banking-ai-mule-detection-confluent\/\" target=\"_blank\" rel=\"noopener\">Real-Time Banking Data And AI-Powered Mule Detection with Confluent Platform<\/a><\/li>\n<li><a href=\"https:\/\/opstree.com\/blog\/data-integration-with-azure-event\/\" target=\"_blank\" rel=\"noopener\">Modernizing Healthcare Data Integration with Azure Event Hubs \u2013 OpsTree<\/a><\/li>\n<\/ul>\n<h2>Related Solutions<\/h2>\n<ul>\n<li><a href=\"https:\/\/opstree.com\/services\/database-and-data-engineering\/\" target=\"_blank\" rel=\"noopener\">Cloud-native data warehouse optimization<\/a><\/li>\n<li><a href=\"https:\/\/opstree.com\/services\/cloud-migration-and-modernization-services\/\" target=\"_blank\" rel=\"noopener\">Cloud security posture management implementation<\/a><\/li>\n<li><a href=\"https:\/\/opstree.com\/services\/devops-and-devsecops-services\/\" target=\"_blank\" rel=\"noopener\">Zero Trust readiness assessment<\/a><\/li>\n<li><a href=\"https:\/\/opstree.com\/services\/cloud-migration-and-modernization-services\/\" target=\"_blank\" rel=\"noopener\">Enterprise Cloud FinOps Implementation Services<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>A practical guide to AI agent observability: how tracing, evaluation and structured debugging help teams monitor quality, investigate failures, manage cost, and operate agentic systems more reliably.\u00a0 Quick answer: AI agent observability is the practice of capturing an agent\u2019s execution path, model calls, retrieval context, tool inputs and outputs, evaluations, latency, token usage and cost [&hellip;]<\/p>\n","protected":false},"author":244582705,"featured_media":32251,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"_coblocks_attr":"","_coblocks_dimensions":"","_coblocks_responsive_height":"","_coblocks_accordion_ie_support":"","jetpack_post_was_ever_published":false,"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_publicize_message":"","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":true,"jetpack_social_options":{"image_generator_settings":{"template":"highway","enabled":false},"version":2}},"categories":[768739552],"tags":[768739732,768739731,768739735,768739710,768739733,768739736,768739734],"class_list":["post-32245","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-development","tag-agent-evaluation","tag-ai-agent-observability","tag-ai-agent-tracing","tag-ai-agent","tag-llm-debugging","tag-opentelemetry","tag-production-ai-monitoring"],"blocksy_meta":[],"jetpack_publicize_connections":[],"acf":[],"jetpack_featured_media_url":"https:\/\/opstree.com\/blog\/wp-content\/uploads\/2026\/09\/AI-Agent-Observability.webp","jetpack_likes_enabled":true,"jetpack_sharing_enabled":true,"jetpack_shortlink":"https:\/\/wp.me\/pfDBOm-8o5","jetpack-related-posts":[],"_links":{"self":[{"href":"https:\/\/opstree.com\/blog\/wp-json\/wp\/v2\/posts\/32245","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/opstree.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/opstree.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/opstree.com\/blog\/wp-json\/wp\/v2\/users\/244582705"}],"replies":[{"embeddable":true,"href":"https:\/\/opstree.com\/blog\/wp-json\/wp\/v2\/comments?post=32245"}],"version-history":[{"count":3,"href":"https:\/\/opstree.com\/blog\/wp-json\/wp\/v2\/posts\/32245\/revisions"}],"predecessor-version":[{"id":32252,"href":"https:\/\/opstree.com\/blog\/wp-json\/wp\/v2\/posts\/32245\/revisions\/32252"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/opstree.com\/blog\/wp-json\/wp\/v2\/media\/32251"}],"wp:attachment":[{"href":"https:\/\/opstree.com\/blog\/wp-json\/wp\/v2\/media?parent=32245"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/opstree.com\/blog\/wp-json\/wp\/v2\/categories?post=32245"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/opstree.com\/blog\/wp-json\/wp\/v2\/tags?post=32245"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}