How to safely connect third-party MCP servers
Review MCP tools, credentials and imported instructions before connecting. Apply authorization and approval rules, with AIVAX configuration examples.
To safely connect third-party MCP servers, verify the operator and endpoint, limit credentials and exposed tools, review everything imported into the model's context, and authorize every call, including reads, against the authenticated user and target resource. Require independent approval for sensitive actions and treat returned content as untrusted input. In AIVAX, also review server instructions and remote skills: both imports are enabled by default.
An MCP connection crosses a trust boundary in both directions. The server supplies text that can influence the model; the model proposes operations against systems holding data, money, or customer state. A tool catalog describes that interface. It does not establish permission to use it.
What should you check before connecting a third-party MCP server?
Before enabling a remote server, review its operator, permissions, imported content and revocation process:
- Operator and destination. Verify the endpoint through the service owner's documentation. Prefer an official server where available; review intermediaries' access to credentials and data.
- Effective permissions. Start with a dedicated, narrowly scoped credential. Identify exactly which records it can read and which operations it can perform. A shared administrative key gives every permitted call that key's reach.
- Imported content. Inspect names, descriptions, schemas, server instructions, skills, and sample results. Reject descriptions that ask the model to reveal secrets or override application policy.
- Enforcement and approval. Decide which layer checks the authenticated actor, target record, arguments, and required human decision. Do not substitute a prompt for that check.
- Change and incident handling. Record the reviewed catalog and configuration, test changes against sandbox data, and define how to revoke access. A remote service can change behavior without changing its URL.
These checks address a concrete attack surface. Microsoft's analysis of MCP prompt injection describes tool poisoning: malicious instructions embedded in tool descriptions that the model uses to choose its next action. Tool results can carry indirect prompt injection too. Discovery is therefore already a trust decision, even before the first tool call.
Which protections belong to the MCP client and server?
The MCP tools specification, revision 2026-07-28, requires servers to validate inputs, enforce access controls, rate-limit invocations, and sanitize outputs. It recommends that clients show tool inputs, request confirmation for sensitive operations, validate results before passing them to the model, implement timeouts, and log usage. These are implementation obligations and recommendations, not proof that a particular connection implements them.
Client controls differ. OpenAI's Responses API MCP integration defaults to approval requests for each MCP call and offers allowed_tools. Anthropic's connector supports allowlisting, denylisting, and per-tool configuration. Do not assume those controls exist in every gateway.
The MCP authorization specification defines an OAuth-based flow for HTTP transports; authorization support itself is optional in the protocol. Where that flow is implemented, the security requirements include token audience validation and resource indicators. A server must reject tokens intended for another resource and must not pass the client's token through to an upstream API. Authentication and delegated scopes still do not decide whether this user may refund this payment.
The MCP security guidance also covers SSRF and scope minimization. If your integration follows URLs from metadata or results, restrict destinations and protect internal services. An approved server can still return an unsafe URL.
Where should each trust-boundary check run?
This diagram shows the enforcement architecture to implement, not security controls automatically supplied by connecting AIVAX to a server.
Every connection forces the same four decisions, and each belongs to a different phase of the request:
| Phase | The question | Where it belongs |
|---|---|---|
| Discovery-time policy | Which tools from this server become visible to this agent, and does the metadata still describe what they do? | Integrator and server: tool selection, inventory, review when behavior changes. |
| Call-time authorization | Whose identity, scopes, and approvals apply to this specific call? | MCP server and application: identity resolution, scope checks, approval rules, idempotency. |
| Result handling | Is this output data or an instruction, and what may re-enter the model context? | Client and server: sanitization, treating content blocks as untrusted input. |
| Audit and replay | Can you reconstruct who did what, approved by whom, and reverse or replay it? | Application and server: authenticated actor, redacted arguments, decision, approver, and operation ID. |
Visibility is not permission. AIVAX caches MCP discovery for 600 seconds by default, so unchanged connection settings can reuse an older catalog after the remote server changes it. That cache does not authorize calls or delay server-side revocation. Enforce permissions on every invocation; do not wait for a catalog refresh to deny access.
How do you authorize a refund tool call?
Consider a support gateway connected to an internal MCP server exposing three tools:
lookup_customer_by_email— reads a customer record; read-only, but still subject to authorization and data minimization.create_support_ticket— opens a ticket; write-capable, low blast radius.issue_refund— moves money; write-capable, high blast radius.
A read-only lookup can still expose another customer's record. Authorize each lookup against the acting user and return only the required fields. Keep ticket creation scoped to support workflows, and expose refunds only through a server or credential whose permissions match the intended workflow.
This illustrative tools/call body uses fictional values and assumes you implemented issue_refund on your own server. AIVAX sends the original remote tool name and adds _meta; the model proposes the arguments. The example is not an executable refund service:
{
"jsonrpc": "2.0",
"id": 41,
"method": "tools/call",
"params": {
"name": "issue_refund",
"arguments": {
"payment_id": "pay_8f21c",
"amount_minor": 4200,
"reason": "duplicate_charge",
"idempotency_key": "refund-pay_8f21c-4200"
},
"_meta": {
"_aiv_external_user_id": "agent-7734",
"_aiv_call_source": "WebChatClient",
"_aiv_conversation_token": "cnv_01J9Q8V2",
"_aiv_nonce": "<BCRYPT_HASH_OF_ACCOUNT_HOOK_KEY>",
"_aiv_moment": "2026-10-05T10:00:00.0000000-03:00"
}
}
}
The server must authenticate the connection, resolve the external ID through a trusted application mapping, check ownership of pay_8f21c, validate the amount, and require a recorded approval before executing. An ID supplied in request context is not proof of the user's identity. A model-generated “approved” argument is not human approval.
_aiv_nonce is a BCrypt hash of the account's configured hook key, or null when no key is configured. The receiver can verify it against that key, but it does not sign the arguments or prevent replay. _aiv_moment is a timestamp, not a freshness guarantee. Use transport authentication and your own replay controls; treat _aiv_conversation_token as correlation, not a unique payment operation.
Likewise, validate and persist the operation's idempotency key on the server. Bind it to the actor and canonical refund parameters; reject conflicting reuse. Do not rely on the model to repeat the same key. Return a concise denial or confirmation and log the authenticated actor, redacted arguments, decision, approval reference, and operation ID. Avoid retaining credentials or full sensitive tool results in audit logs.
How do you configure a third-party MCP source in AIVAX?
AIVAX's gateway MCP client uses Streamable HTTP and configured custom headers. It lists remote tools and exposes their descriptions and input schemas as model-callable functions, with names prefixed by the source. The source configuration has no dedicated allowed_tools field: limit the catalog on the server or through a reviewed intermediary. Generic skill-based tool visibility settings are not server-side authorization.
The following is a gateway configuration fragment, not a complete inference request. Replace the example endpoint and token with values for a server you control or have reviewed. Keep the credential in trusted configuration, never in a prompt or frontend bundle. When updating an existing gateway, merge deliberately: assigning parameters.mcpSources replaces that array, rather than appending one source.
{
"parameters": {
"mcpSources": [
{
"name": "support",
"url": "https://mcp.example.com/mcp",
"headers": {
"Authorization": "Bearer <MCP_SERVER_TOKEN>"
},
"cacheDuration": 600,
"allowClientInstructions": false,
"allowRemoteSkills": false
}
]
}
}
This disables two optional imports; it does not filter tools or make their content safe. See the MCP configuration reference and AI Gateway guide for the surrounding contract.
What crosses into context depends on those settings:
- Tools and results: descriptions and schemas come from the server. Text and structured results return to the model; supported image and audio blocks can also enter the conversation. Parsing a result does not sanitize prompt injection.
- Server instructions:
allowClientInstructionsdefaults totrue; advertised usage instructions are added to the gateway's system instructions. Disable this unless you intend to trust that instruction source. - Remote skills:
allowRemoteSkillsdefaults totrue. Supported rootSKILL.mddocuments are imported after size, digest, and metadata checks. The model can load their content as skills. Matching a server-provided digest proves consistency with its manifest, not independent trustworthiness. Supporting files are not imported; see MCP skills and workflow boundaries.
Discovery caches tools, instructions, and skills, not individual call results. Its configured duration defaults to 600 seconds. The gateway MCP connection uses supplied headers; it does not perform OAuth consent or token refresh on your behalf. Supplying a bearer token is not the same as managing its lifecycle.
AIVAX also has gateway workers: application-owned hooks can block ordinary server-side tool calls before execution. They are not a ready-made human approval workflow. Check the actual execution path before relying on them; a tool wrapped inside a shell command does not necessarily produce its own worker event. Keep authorization on the receiving MCP server regardless.
What should you test before giving the connection production access?
Run these checks with test credentials and sandbox records, not real payments:
- Cross-user access: try another user's customer or payment ID. The server must deny it even when the JSON matches the tool schema.
- Unapproved writes: request a refund without an independently recorded approval. Confirm that no external write occurred; the assistant's answer is not evidence.
- Duplicate delivery: repeat an approved operation. Verify a single side effect and reject reuse of its key with changed parameters.
- Injected content: return a harmless test instruction asking the model to ignore policy. Verify that a subsequent unauthorized call remains blocked outside the model, even if the model follows the text.
- Revocation and change: revoke the credential or user permission while discovery is cached. Calls must fail at the server. Review changed descriptions, instructions, and skills before accepting the new catalog.
State handles need the same ownership checks. The companion guide to MCP's stateless core and application-owned state explains why a conversation or resource identifier does not replace authorization.
Cutting a server's catalog to fewer tools also shrinks what the model can be told to call, but it does not replace these checks; the guide to reducing MCP tool-definition tokens covers the options and their limits.
Review data retention separately for each processor. OpenAI's guide says that, with store=true in the Responses API, data sent to MCP servers is logged for 30 days unless Zero Data Retention is enabled. The third-party server's own retention and residency policies still apply. Anthropic's MCP connector documentation says it is not covered by ZDR arrangements. These are those integrations' terms, not AIVAX retention guarantees.
Frequently asked questions
Are read-only MCP tools safe to enable without approval?
Not automatically. A read can disclose another customer's data, and the query arguments themselves can send sensitive information to a third party. Returned text can also influence later calls. Reduce privileges, enforce record-level access, and decide approval requirements from the data exposure—not just whether the tool writes.
Does OAuth prevent MCP prompt injection or tool poisoning?
No. OAuth controls delegated access; it does not make tool descriptions, instructions, or results trustworthy. A malicious response can persuade a model to propose an action that its token permits. Combine scoped credentials with content review, independent authorization, and approval for sensitive operations.
Does AIVAX ask a human before every MCP tool call?
No built-in per-call human approval flow is provided by the gateway MCP source. Implement approval in your application and enforce it where the action executes. Workers can contribute policy checks, but a model confirmation, source visibility setting, or valid argument schema is not an approval.