- Writing status
- Finished
- Assumed audience
- Architects deciding whether to build an API or an MCP server, and where authorization should live once they have both.
- Key takeaways
- An API and an MCP server answer different questions and compose as layers; asking which to build is the wrong question.
- Authentication, authorization and business policy belong in the HTTP API, implemented exactly once.
- An MCP server should be a thin translation layer that contains none of that policy.
If you run a service that AI agents should be able to operate, you will face a design question that the current tooling discussion frames badly: should you build an API or an MCP server? This article argues, from two production implementations, that the question is malformed. An API and a Model Context Protocol server answer different questions, they compose as layers, and the design decision that actually matters is where the policy lives. The rule this article defends: all authentication, authorization, and business policy belongs in the HTTP API, implemented exactly once, and any MCP server should be a thin translation layer that contains none of it. I will define both terms, give the rule's rationale, show it running, and state where it might not apply.
What an API is, what MCP is: definitions#
An API, in the sense used here, is a durable HTTP contract: an endpoint that accepts authenticated requests, applies rules, and performs operations. The relevant example on this site is a publishing API: present a bearer token, submit an operation such as saving a post, and the server validates the content, enforces the publication policy, applies rate limits, and lands an atomic git commit. Any HTTP-speaking caller with the credential can use it: a script, a scheduled job, a CI step, or an AI agent.
The Model Context Protocol, introduced by Anthropic in late 2024 and now supported across the major AI vendors, addresses a different problem: discovery and calling conventions for AI assistants. An MCP server describes its tools over the protocol, names, typed parameters, documentation, and a connected assistant can call them without anyone writing integration code specific to that assistant. The description is the integration. This solved a real combinatorial problem, many models times many tools, and it is why so many services became agent-callable so quickly.
Notice the division: the API is capability with rules; MCP is discoverability and calling convention. Neither replaces the other, and the failure mode worth an article is treating them as peers.
The rule: authentication, authorization, and business policy live in the API#
State the rule concretely: when an agent calls an MCP tool on my site, the MCP server translates that call into the same bearer-authenticated HTTP request any other caller would make, and the API decides. The MCP layer holds no validation logic, no permission checks, no rate limiting of the underlying operations, and no knowledge of the publication policy. A check script in its repository asserts this mechanically: no imports from the application codebase, no database binding, no repository credential. If the MCP layer were deleted, the set of allowed operations would not change.
The rationale has three parts, in decreasing order of importance.
First, duplication produces drift. If the MCP layer implemented its own copy of the rules, the system would have two policies that began identical and diverge, because every future change lands in one place first and sometimes never reaches the second. This is not speculative on this site: the same project earlier adopted a one-renderer rule for its content pipeline, after observing that two renderers made it impossible to distinguish content drift from implementation difference. Two enforcement points are the same defect in a different subsystem. A policy that exists in two places is two policies.
Second, auditability. When rules exist exactly once, there is exactly one code path to test, one to review after an incident, and one that can be wrong. The publication policy on this site (an agent may edit and republish but may not perform a post's first publication, covered in full in the trust model article) is enforced in one function, exercised by one test suite, and produces one refusal message. Every caller, human tooling or agent, receives that same refusal verbatim.
Third, stability layering. HTTP with bearer authentication has been stable for decades. MCP is young and moving: its 2026-07-28 specification revision, current as this publishes, is a breaking change, the largest since the protocol launched, removing the session handshake and reworking authorization, though it arrives with a formal deprecation lifecycle promising twelve-month windows in the future. None of that is a criticism; it is what healthy young protocols do. It is, however, a strong argument about ordering: volatile layers belong on top of durable ones. When this site's MCP layer needed rework to track the new revision, the rework touched translation only. The policy did not move, because the policy does not live there.
Every kind of author enters through the same door
The MCP server is one more caller, not a second entrance. Delete it and the set of allowed operations is unchanged, which is the property the rule buys.
The same layering on the read side: three presentations, one engine#
The publish path was not the first place this site used the pattern. Its search engine ships at three levels: a plain URL anyone can construct, the same URL returning JSON under HTTP content negotiation, and an MCP endpoint an assistant can query conversationally. Three presentations, one engine. The removability of the top layer was verified by removal: with the MCP level off, the classic search response was byte-identical to before it existed. A presentation layer you can remove without touching behavior is a presentation layer wired correctly, and the test is cheap enough to run rather than assert.
When to use the API directly and when to use MCP#
Use the API directly when the caller is software you control or software that should outlive protocol churn: scripts, cron, deployment tooling, tests, integrations you write yourself. Use MCP when the caller is an AI assistant and the value is ambient availability: tools that appear in a conversation, described well enough to be used correctly without bespoke glue.
Two craft points for the MCP layer, both cheap and both frequently skipped. Write the policy into the tool descriptions, so an agent learns the rules before its first call; on this site, the save tool's description states the first-publication restriction and says explicitly that the refusal is correct behavior rather than an error to retry. And pass the API's error messages through verbatim rather than summarizing them, because the API's refusals name their policy and the permitted alternative, and a translation layer that paraphrases the lock misinforms the visitor.
One structural asymmetry is worth designing in deliberately: anything the MCP layer can do, the API can do, and not the reverse. If you find an operation possible through your MCP server that is not possible through your API, policy has leaked into the presentation layer, and the audit story from the rationale above no longer holds.
Where the rule might not apply#
The argument above is scoped to a particular situation: a service with meaningful rules, operated by identified callers, where policy drift and unaudited writes are the expensive failures. Different situations weigh differently, and honesty requires naming a few. A read-only MCP server over public data has little policy to misplace, and building it standalone is fine. A team standardized on an MCP-native gateway with centralized authorization may reasonably put enforcement in that gateway, which then simply is their API in this article's sense, wearing a different protocol. And MCP's own authorization story matured substantially in the 2026-07-28 revision, formalizing servers as OAuth 2.1 resource servers; identity of the caller can and should live at the MCP layer, which is distinct from policy about operations. On this site, those are literally two credentials: an OAuth flow answers who is operating, and the API's bearer token governs what operators may do. Keeping the two questions separate is what lets an authentication failure and a policy refusal read differently, because they are different.
The compressed form of the recommendation, for architects skimming: build the door once, with the lock in it, and add doorbells freely, provided every doorbell is only a doorbell. This is the seventh post in the series, following the AI answer layer. The next article examines the lock itself: the trust model for giving an AI agent write access to a production system, and the one operation it is structurally prevented from performing; the final one covers building the MCP server under the new specification.
Update, August 2026#
The rule got a month of adversarial testing it did not have at publication. Two external audits of this site went looking for policy in the wrong layer and found none there, and every hardening the audits prompted landed where the rule says it must: a rate limit on the authentication callback, pacing on the answer layer's daily ceiling, and a compensation path for the one write-failure window all went into the API, and the MCP layer's diff for the whole month is empty. That is the drift argument running in reverse and it is the cheapest audit story I have ever had: when the policy can only be one place, the review of a security month is a review of one directory. One caution earned rather than reasoned: the door metaphor extends to doors you forgot you have. The audits' three real findings were all on HTTP routes that predated the rule and had never been walked through it, an unauthenticated delete among them. The rule only protects the operations you route through the door, and the standing obligation it creates is the inventory, the same lesson the trust model article reports from the visibility side.
Built on the stack described at /colophon.