MCP permission design is the part most teams skip until an agent can read too much, write too much, or do both too quickly. If you are putting tool-using AI into a real business workflow, the permission model is the architecture.
That is where trust lives. Not in the model. Not in the prompt. In the boundaries around what the agent can see, call, and change.
- MCP Permission Design Basics
- Scope Tools by Job, Not by User
- Approval Flows and Human Signoff
- Audit Logs and Replayable Decisions
- Failure Modes and Hard Limits
MCP Permission Design Basics
The first mistake is treating MCP like a generic API wrapper. It is not. An MCP server is a tool boundary, and tool boundaries need explicit permission design. If your agent can invoke a calendar tool, a CRM write tool, and a billing export tool, those are three different risk profiles. They should not share one broad token with one broad scope.
The cleanest model is simple: read, propose, and act. Read tools can inspect state. Propose tools can prepare a change set. Act tools can mutate state, but only after a policy check. That policy should not live in the prompt. It should live in the server or an authorization layer sitting in front of it. We use this kind of separation when we build internal AI workflows in MCP Server Architecture: Securing LLM Tool Execution.
Think about concrete examples. A support agent can query ticket history. Fine. It can draft a refund. Fine. It should not issue the refund until a human approves it or a policy engine confirms the amount is under threshold, the account age is above threshold, and the refund count is below threshold. That is not being cautious for its own sake. That is how you keep an autonomous workflow from becoming an incident.
A practical starting point looks like this:
{
"tool": "issue_refund",
"scope": "payments.refunds.write",
"max_amount_cents": 5000,
"requires_approval": true,
"allowed_context": ["support_case_id", "customer_id"]
}
Notice what is missing: free-form instructions, natural language exceptions, and hidden assumptions. Good MCP permission design is boring on purpose. It is a table of allowed actions, the preconditions for each, and the failure path when a condition is not met.
Scope Tools by Job, Not by User
Most permission systems are built around humans. MCP permissions should be built around jobs. A person may be a support manager, but the agent running on their behalf might only be allowed to search tickets, summarize history, and open a refund draft. The agent does not need the same authority that the human has in the admin console. That distinction matters.
This is where role-based access control starts to fray if you are not careful. RBAC is fine for coarse boundaries, but agents need task-scoped permissions. A job to “investigate a failed onboarding” may need access to auth logs, recent webhook events, and one customer record. It does not need export access to the entire tenant, even if the human in the loop has that ability. You are minimizing the blast radius of a mistaken tool call.
For a real implementation, I prefer a permission matrix with three dimensions: tool, resource, and action. For example, an MCP server for internal ops might allow:
- read: customer profile, order history, incident timeline
- propose: refund, password reset, subscription downgrade
- act: password reset only for accounts under a threshold and only with approval
That matrix maps cleanly onto policy engines like OPA or Cedar, and it fits well with a narrow service layer in Node.js or Go. If you already have internal service boundaries, this is the same pattern you use for humans, just with stricter defaults. The agent is a caller. It is not a special class of principal.
One useful rule: if a tool would be dangerous in a wrong tab, wrong tenant, or wrong hour, it needs another gate. That may be tenant scoping, time-based restrictions, or environment-based restrictions. A finance agent should not be able to write to production billing at 2 a.m. just because the prompt sounded confident.
Approval Flows and Human Signoff
Approval is not a checkbox. It is a workflow. If you design it badly, humans rubber-stamp the agent. If you design it well, the agent does the tedious parts and the human signs off on the irreversible parts. That is the difference between assistance and delegation.
A good approval flow has three states: draft, review, and commit. The agent can prepare a draft action with all the evidence attached. The reviewer sees the exact tool call, the input parameters, the policy checks that passed, and the diff or side effect that will occur. Only then does the system commit. For high-risk actions, the commit token should be short-lived and single-use.
Here is the failure mode I see most often: teams let the agent re-run its own approved action with slightly different parameters after a human has already signed off. That creates a gap between intent and execution. The fix is to bind approval to a hash of the proposed action. If the parameters change, approval expires. The human approved one exact thing, not “something similar.”
One pattern that works well is a two-channel UI: the agent chat on the left, the approval card on the right. The approval card should show a structured summary, not prose. For example:
- Action: disable stale API key
- Principal: svc-agent-ops
- Resource: key_7f3a
- Risk: medium
- Reason: inactive 120 days, last use from deprecated app
- Required approver: on-call engineer
This is the same kind of discipline you see in CI/CD Security Hardening: Protecting Your Pipeline. The human is not there to read every log line. The human is there to confirm the boundary was respected. If you keep the approval packet small, clear, and deterministic, reviewers actually use it.
Audit Logs and Replayable Decisions
If you cannot replay a decision, you do not really understand it. MCP permission design needs audit logs that capture the prompt context, the tool selection, the policy evaluation, the human approval, and the exact response from the tool. Not a sentence summary. The actual record.
This matters for security, but it also matters for debugging. When an agent does the wrong thing, the root cause is often one of four things: a bad prompt, a bad tool schema, a missing policy rule, or stale state. A replayable log lets you isolate which one failed. Without that, you are guessing from screenshots and Slack messages.
Use structured events. A JSON log record should include correlation IDs, tenant IDs, tool name, policy version, and a decision outcome. For example:
{
"trace_id": "9c1a...",
"tenant_id": "acme-42",
"tool": "create_case_note",
"policy_version": "2026-01-14",
"decision": "allow",
"approval_id": "apr_8831",
"input_hash": "sha256:...",
"result": "ok"
}
That log becomes your control plane. It is what security reviews, incident reviews, and SOC 2 evidence collection all want, even if they ask for it in different words. If your org is already thinking about evidence, the discipline in SOC 2 Evidence Automation for Engineering Teams maps cleanly here.
Replayability also lets you run offline evaluation. Take 100 real tool decisions, scrub sensitive fields, and replay them against a new policy version before you ship it. That catches regressions in allow/deny behavior before users do. It is the same idea as testing database migrations against a production clone: trust the change less than the tool vendor tells you to.
Failure Modes and Hard Limits
The hard truth is that every agentic system fails at the edges. MCP permission design should make those edges obvious. If the model is uncertain, the server should deny by default. If the tool schema is incomplete, the call should fail closed. If the tenant context is missing, the request should never reach execution.
There are a few failure modes worth naming explicitly. First, scope creep: a tool starts narrow and accumulates flags until it becomes a back door. Second, approval fatigue: humans approve everything because the system asks too often. Third, context bleed: the agent carries data from one tenant or case into another. Fourth, silent retries: the agent repeats a failed action until something bad happens. These are not theoretical. They show up in the first serious deployment.
The fix is a hard limit stack. Set max tool calls per task. Set max mutation count per session. Block cross-tenant access at the service layer, not only in the prompt. Require explicit re-authorization when the agent changes from read to write mode. And for tools that can move money, change identity, or expose secrets, keep a human approval path even if the model looks “confident.” Confidence is not authorization.
If you want a quick decision matrix, use this:
- Read-only data: agent can call directly with tenant-scoped token
- Low-risk write: agent can draft, system auto-commits under policy thresholds
- High-risk write: agent drafts, human approves, system commits once
- Secrets or identity: agent never sees the secret, only the outcome
That last line is the one teams resist. They want the agent to “just do it.” That is how incidents get written up later. The better path is narrower than people like, but it lasts longer.
Bad permission design turns an AI workflow into an uncontrolled operator. If you are mapping real tool access across tenants, approvals, and audit trails, it is worth treating the permission layer as first-class architecture. If that is the problem in front of you, you can apply for an engagement; the application takes ten minutes. We take three engagements a quarter, and Sprint work is built around one shipped outcome.




