All articles

English

September 22, 20265 min read

MCP in Production: Architecture, Security, and When Not to Use It

A practical guide to production MCP architecture: transports, trust boundaries, authorization, observability, and the cases where a direct API remains the better choice.

Portrait of Tran Kim Dat

Tran Kim Dat

Full-stack Engineer

Abstract production architecture with a central routing node connecting secured clients and isolated servers

Model Context Protocol has moved from an interesting interoperability idea to a serious production boundary. The important question is no longer whether an agent can connect to a tool. It is whether that connection preserves identity, limits authority, survives failure, and leaves enough evidence to explain what happened.

This guide treats MCP as infrastructure, not magic. It explains the host–client–server architecture, the July 2026 protocol changes, the security controls that matter, and the situations where a direct API call is still the cleaner engineering choice.

What MCP standardizes—and what it does not

MCP gives AI applications a common way to discover and invoke server capabilities. Servers can expose tools for actions, resources for contextual data, and prompts for reusable interaction patterns. A host application creates a client for each server connection and remains responsible for orchestration, permissions, consent, and model access.

That separation is useful because integrations stop being hard-coded to a single agent or model provider. But MCP does not make an unsafe tool safe. It does not decide who should be allowed to delete a database row, which data may enter a model context, or when a human must approve an action. Those are product and security responsibilities around the protocol.

Diagram of an MCP host, client, policy boundary, and servers exposing tools, resources, and prompts
A production MCP deployment is a trust architecture: the host owns user intent, the client maintains protocol connections, and policy controls access to server capabilities.

The 2026 architecture is intentionally more web-native

The final 2026-07-28 MCP specification removed the mandatory initialization handshake and protocol-level session requirement. Requests can be stateless, cacheable, and easier to route through conventional web infrastructure. The release also added clearer HTTP headers, stronger authentication guidance, W3C trace-context support, and cache controls for tool metadata.

For local development, stdio remains attractive: the host launches a child process and communicates over standard input and output. It is simple and keeps the server off the network. For shared or remote production services, Streamable HTTP is the natural fit because it can sit behind normal authentication, routing, rate limiting, and observability layers.

Choose the transport from the trust boundary outward. “Local versus remote” is a security and operations decision before it is a protocol decision.

A production-ready request path

A robust request path has six distinct responsibilities:

  1. Authenticate the person and workload. Establish both user identity and the calling application or agent identity.

  2. Authorize the capability. Filter discovery and invocation so an agent sees only the tools required for the task.

  3. Validate arguments. Treat model-produced parameters as untrusted input, even when they match a declared schema.

  4. Apply execution policy. Add rate, cost, environment, tenant, and data-sensitivity limits outside the model.

  5. Request approval when impact is high. Writes, financial actions, privilege changes, and destructive operations deserve explicit checkpoints.

  6. Record the outcome. Correlate the user request, model decision, tool call, downstream response, and final answer.

Tool descriptions should describe intent, inputs, outputs, and side effects. They should not embed secrets or trust data returned by external systems as instructions. A production server also needs timeouts, bounded response sizes, idempotency where applicable, and a stable error vocabulary.

Least privilege starts at discovery

If a client can discover two hundred tools, the model has a larger decision surface and the system has a larger attack surface. Prefer task-shaped capabilities over raw administrative APIs. A tool named prepare_invoice_preview can validate business rules and return a dry run; a generic execute_sql tool pushes too much authority into probabilistic planning.

Recent daily.dev discussions make the same operational point from several angles: production teams are adding gateway-level tool provisioning, short-lived credentials, just-in-time approval, and audit trails. Those patterns are useful signals from practitioners, but the durable principle is simpler: policy must be enforced by code the agent cannot rewrite.

Prompt injection crosses tool boundaries

An MCP server may retrieve tickets, documents, web pages, or repository content. Any of that data can contain instructions designed to redirect the agent. Keep trusted system guidance separate from untrusted tool output. Label provenance, sanitize rendered content, constrain follow-up actions, and never let a retrieved sentence silently expand permissions.

Authentication alone does not solve this. A fully authenticated agent can still be manipulated into making an authorized but unwanted call. The defense is layered: narrow tools, contextual authorization, approval gates, output validation, and monitoring for unusual sequences.

Observability should follow the full chain

Trace propagation in the 2026 specification makes it easier to connect an MCP request with application and downstream service traces. Use a correlation identifier across the host, client, gateway, MCP server, and the API the tool eventually calls. Capture duration, result class, retry count, token usage, and policy decisions. Redact secrets and sensitive payloads by default.

Logs answer what happened; traces explain where time and failure accumulated; evaluations answer whether the final behavior was good. You need all three for an agent that can call real systems.

When MCP is the wrong abstraction

Do not build an MCP server merely because the protocol is popular. A direct API or typed SDK is usually better when there is one application, one stable integration, a small stateless operation, and no need for cross-client discovery. An unnecessary MCP wrapper creates a second contract that can drift from the underlying API.

MCP earns its place when it provides at least one of four things:

  • Orchestration: one safe, named capability replaces a fragile multi-call sequence.

  • Framing: a large API is reduced to the small surface an agent should use.

  • Reach: the server adapts a legacy, local, or non-HTTP system.

  • Control: the boundary adds authorization, approvals, redaction, or auditability.

This “MCP versus curl” test, also debated on daily.dev, is a useful antidote to protocol-driven design. If the wrapper adds none of those properties, keep the simpler path.

A practical launch checklist

  • Document the host, client, server, and downstream trust boundaries.

  • Use short-lived, audience-bound credentials and avoid passing provider secrets through prompts.

  • Filter tools per user, tenant, environment, and task.

  • Separate read, preview, commit, and destructive actions.

  • Add timeouts, payload limits, rate limits, and idempotency keys.

  • Propagate traces and retain policy decisions without storing sensitive content unnecessarily.

  • Test prompt injection, stale tool metadata, partial failure, retry behavior, and revoked access.

The decision

MCP is valuable because it standardizes a boundary, not because it eliminates architecture. In production, the winning design keeps the model inside a narrow loop: discover only what is relevant, prove identity, evaluate policy, ask for approval when impact rises, execute a bounded action, and record the result. If a direct API already does that with less machinery, use the direct API.

Primary sources and further reading

Information checked on September 22, 2026.

Continue reading

More field notes.

View all articles
Application traffic flowing through a policy and telemetry gateway toward several model providers
EnglishSep 22, 2026

AI Gateway vs Direct Provider SDKs in a Production Next.js App

A practical architecture comparison for Next.js teams choosing between direct model-provider SDKs and an AI gateway for routing, failover, policy, and observability.

ai-gateway · ai-sdk · architecture · llm · nextjs

Read article
Three abstract model cores of different sizes connected to code, tools, memory, and infrastructure symbols
EnglishSep 22, 2026

Open-Weight Coding Models in 2026: A Practical Selection Guide

How to evaluate open-weight coding models by task quality, hardware, context, tool use, license, privacy, and operating cost—not benchmark headlines alone.

AI · coding-models · developer-tools · llm · open-source

Read article
Abstract distributed trace linking model, tool, retrieval, and analytics signals in a dark observability system
EnglishSep 22, 2026

Open-Source Agent Observability with OpenTelemetry and Langfuse

Instrument an AI agent end to end with OpenTelemetry semantics and Langfuse: traces, model and tool spans, evaluations, privacy controls, and useful alerts.

AI Agents · langfuse · observability · open-source · opentelemetry

Read article

Have a product worth building carefully?

Start a conversation