All articles

English

September 22, 20265 min read

Reliable AI Agents: State, Tool Budgets, Guardrails, and Human Review

A production-minded blueprint for AI agents that can recover state, limit tool use, enforce guardrails, escalate risky actions, and stop predictably.

Portrait of Tran Kim Dat

Tran Kim Dat

Full-stack Engineer

Abstract agent loop enclosed by safety barriers, review controls, and a finite resource meter

An AI agent is not reliable because it chose a strong model. It is reliable because the surrounding system constrains what the model can do, persists what must survive, and knows when to stop.

That surrounding system is often called an agent harness: tools, context assembly, state, policies, budgets, observability, evaluation, and human control. The model proposes the next step; the harness decides whether that step is valid, affordable, authorized, and recoverable.

Start with the simplest control flow

Anthropic’s guidance separates workflows, where code controls a known sequence, from agents, where the model chooses its own process and tool use. Use a workflow when the path is predictable. Add agentic freedom only where the task genuinely needs open-ended planning.

This is not a philosophical distinction. Deterministic code is easier to test, cheaper to run, and simpler to recover. A useful production system may be mostly workflow with one agentic decision point.

Reliable agent loop showing plan, act, observe, and check stages with budgets, guardrails, and exit conditions
Reliability comes from the harness around the model: explicit state, bounded actions, policy checks, and more than one safe way to stop.

State is a product contract

Chat history is not enough. A long-running agent needs explicit state: goal, current plan, completed steps, tool results, approvals, remaining budgets, retry counts, and the reason it stopped. Store durable state outside the prompt so an interrupted run can resume without asking the model to reconstruct reality from a transcript.

Separate four kinds of state:

  • Conversation state for what the user said and what the system answered.

  • Workflow state for steps, dependencies, checkpoints, and status.

  • Domain state for records in the product itself.

  • Execution evidence for tool inputs, outputs, policy decisions, and trace identifiers.

Make every write idempotent where practical. A retry after a timeout should not create a second invoice, send the same email twice, or approve the same deployment again.

Budgets turn an open loop into a bounded system

Every agent run should have limits the model cannot increase: maximum steps, tool calls, elapsed time, input and output tokens, monetary spend, concurrent operations, and retries per dependency. The correct limit depends on the task, but the absence of a limit is itself a design defect.

Track budgets in the orchestrator, not in natural-language instructions. Before each action, reserve the expected cost; after the action, record actual usage. When a budget is nearly exhausted, the agent can summarize progress, ask for a decision, or return a partial result instead of failing silently.

Guardrails belong at several layers

Input guardrails can reject malicious or out-of-scope requests before expensive work begins. Tool guardrails can validate arguments and outputs around each action. Output guardrails can check the final answer for policy, privacy, or schema requirements.

Ordering matters. The OpenAI Agents SDK documentation notes that a parallel input guardrail may finish only after the model has already consumed tokens or invoked tools. For high-risk requests, run the guardrail in blocking mode or enforce the same policy at the tool boundary. A guardrail that arrives late is monitoring, not prevention.

Use deterministic checks for deterministic rules: schemas, allowlists, tenant ownership, dollar ceilings, path restrictions, and forbidden operations. Use model-based checks for semantic judgments, then define what happens when that checker is uncertain.

Human review should be designed, not improvised

Human-in-the-loop is most effective when the reviewer receives a compact decision package: proposed action, target, material inputs, expected effect, reversible plan, evidence, and the exact permission being requested. “Approve?” without context merely transfers confusion to a person.

Create explicit risk tiers. Read-only retrieval may proceed automatically. Draft generation can proceed but not send. External communication, data mutation, payments, privilege changes, and destructive actions can require approval. Very high-risk actions may be prohibited entirely.

Approval should authorize a specific action with an expiry, not give the agent broad standing permission. Revalidate policy after approval because the underlying resource or parameters may have changed.

Failure is a normal state

Agents face ordinary distributed-system failures plus uncertain model behavior. Tool calls time out. Providers return rate limits. Context becomes stale. A model repeats itself. A human never responds. Define transitions for each condition:

  • Retry only transient, idempotent work with backoff and a hard cap.

  • Replan when the tool result changes the available path.

  • Pause when user input or approval is required.

  • Degrade to a smaller capability when a dependency is unavailable.

  • Stop on policy violation, exhausted budget, repeated failure, or loss of authorization.

Observe decisions, not private reasoning

Record the inputs and outputs that explain system behavior: selected tool, validated arguments, response class, latency, token usage, policy result, approval result, and final outcome. Do not depend on hidden chain-of-thought. A concise decision summary and structured events are more useful for debugging and safer to retain.

Tracing should connect model calls, retrieval, tools, and application code. Evals should measure task success, policy compliance, tool selection, citation quality, and recovery behavior. Operational dashboards then watch latency, cost, error rate, loop length, approval frequency, and stop reasons.

A minimal production blueprint

  1. Validate the request and assign a risk tier.

  2. Load only the context and tools allowed for that user and task.

  3. Create a durable run record with fixed budgets.

  4. Let the model propose one next action.

  5. Validate policy and arguments before execution.

  6. Execute through an idempotent adapter and capture the result.

  7. Evaluate progress against explicit completion and stop conditions.

  8. Finish, pause for review, replan, or stop safely.

Recent daily.dev coverage of production agents repeatedly surfaces the same checklist—model control, prompt versioning, guardrails, budget limits, tool authentication, tracing, and evals. The technologies will change; these control points are the stable architecture.

The practical rule

Give the model freedom over decisions that benefit from judgment, and keep authority over everything that must remain predictable. Reliable agents are not unconstrained autonomous workers. They are bounded systems with explicit state, narrow tools, enforceable policy, measurable behavior, and graceful exits.

Primary sources and further reading

Information checked on September 22, 2026.

Continue reading

More field notes.

View all articles
Application traffic flowing through a policy and telemetry gateway toward several model providers
EnglishSep 22, 2026

AI Gateway vs Direct Provider SDKs in a Production Next.js App

A practical architecture comparison for Next.js teams choosing between direct model-provider SDKs and an AI gateway for routing, failover, policy, and observability.

ai-gateway · ai-sdk · architecture · llm · nextjs

Read article
Three abstract model cores of different sizes connected to code, tools, memory, and infrastructure symbols
EnglishSep 22, 2026

Open-Weight Coding Models in 2026: A Practical Selection Guide

How to evaluate open-weight coding models by task quality, hardware, context, tool use, license, privacy, and operating cost—not benchmark headlines alone.

AI · coding-models · developer-tools · llm · open-source

Read article
Abstract distributed trace linking model, tool, retrieval, and analytics signals in a dark observability system
EnglishSep 22, 2026

Open-Source Agent Observability with OpenTelemetry and Langfuse

Instrument an AI agent end to end with OpenTelemetry semantics and Langfuse: traces, model and tool spans, evaluations, privacy controls, and useful alerts.

AI Agents · langfuse · observability · open-source · opentelemetry

Read article

Have a product worth building carefully?

Start a conversation