Token Vault
Reference

MCP Endpoint

Reference for POST /api/agents/mcp — the MCP server exposing list_credentials and get_credential tools, with tvagent_ key or OAuth 2.1 session auth.

POST /api/agents/mcp

Token Vault's own MCP server. Any MCP client (Claude Code, Claude.ai, ChatGPT, Cursor) can connect over streamable HTTP and use standard tool calls to work with the agent's granted credentials.

Claude Code
claude mcp add --transport http token-vault https://api.tokenvault.uk/api/agents/mcp

Authentication

Two interchangeable identities — both resolve to the same agent, grants, and policies:

  • OAuth 2.1 (recommended for MCP clients) — the client discovers Token Vault's authorization server from a 401, registers itself, and obtains a tvsess_… session token in-client. No key is pasted anywhere. See OAuth 2.1 for MCP Clients.
  • Static key — Authorization: Bearer tvagent_… for server-side agents you configure by hand.

Auth is required on every JSON-RPC method, including initialize. An unauthenticated (or invalid) call to initialize gets the same 401 + WWW-Authenticate discovery challenge as any other method:

WWW-Authenticate: Bearer resource_metadata="https://api.tokenvault.uk/.well-known/oauth-protected-resource"

In practice this changes nothing for a real client — the same header you already send on tools/call is just checked one method earlier. It exists because this endpoint shares its rate-limit/penalty-box state with the REST credentials endpoint, which needs a resolved identity to key on before any of the layers below can run.

Tools

ToolArgumentsReturns
list_credentials—The agent's active grants, enriched with service metadata (description, display name, tags, token type, expiry, whether it carries a refresh token) — no credential material
get_credentialserviceA signed, short-lived credential URL pointing at your webhook's /v1/credential
whoami—This agent's own identity: id, name, status, owner, active grants, and the names of any ABAC policies attached to it — no credential material
show_credentialserviceMetadata for ONE granted service — the same fields a list_credentials entry carries — without retrieving the credential itself
credential_historyservice, limit (optional, max 50), cursor (optional)This agent's own recent access history for one granted service: event type, timestamp, and outcome — no credential material

Why a URL and not the credential?

The MCP transport runs through Token Vault, and Token Vault never carries credential bytes. get_credential therefore returns a one-time signed URL; the agent fetches it and receives the credential directly from your webhook — the same custody path as the REST endpoint's 307 redirect. Every other tool on this page — list_credentials, whoami, show_credential, credential_history — only ever returns metadata Token Vault already holds in its own control plane (grants, agent identity, activity log). None of it is the credential, derived from the credential, or a substitute for fetching it.

list_credentials

No arguments. Each granted service is enriched with Firestore-owned metadata (description, displayName, tags) and a whitelisted slice of your webhook's own list response (tokenType, expiryTime, hasRefreshToken). The webhook lookup is best-effort — if it fails, the result silently falls back to the metadata-only fields instead of erroring.

tools/call result — structuredContent
{
  "grants": [
    {
      "serviceName": "github",
      "grantExpiresAt": null,
      "refreshPolicy": "none",
      "description": "GitHub PAT for CI",
      "displayName": "GitHub",
      "tags": ["ci"],
      "tokenType": "pat",
      "expiryTime": null,
      "hasRefreshToken": false
    }
  ]
}

whoami

inputSchema
{ "type": "object", "properties": {}, "required": [] }

Returns this agent's own id, name, status, owner, active grants, and attached policy names — never the owner's email, and never a policy's rule configuration, only its name and rule type(s).

tools/call result — structuredContent
{
  "id": "agent-1",
  "name": "CI Bot",
  "status": "active",
  "ownerId": "firebase-uid",
  "grants": [
    { "serviceName": "github", "grantExpiresAt": null, "refreshPolicy": "none" }
  ],
  "policies": [
    { "name": "Business Hours", "ruleTypes": ["time_window"] }
  ]
}

show_credential

inputSchema
{
  "type": "object",
  "properties": {
    "service": { "type": "string", "description": "The granted service name to show metadata for." }
  },
  "required": ["service"]
}

The same fields a list_credentials entry carries, for one service — metadata only, never the credential. An ungranted service returns the identical, deliberately ambiguous NO_GRANT error get_credential uses, so this can't be used to enumerate services the agent was never granted.

tools/call result — structuredContent
{
  "serviceName": "github",
  "grantExpiresAt": null,
  "refreshPolicy": "none",
  "displayName": "GitHub",
  "tokenType": "pat",
  "hasRefreshToken": false
}

credential_history

inputSchema
{
  "type": "object",
  "properties": {
    "service": { "type": "string", "description": "The granted service name to show access history for." },
    "limit": {
      "type": "integer",
      "description": "Maximum number of history rows to return per page (default 20, max 50).",
      "minimum": 1,
      "maximum": 50
    },
    "cursor": {
      "type": "string",
      "description": "Opaque pagination cursor from a previous call's nextCursor, to fetch the next page."
    }
  },
  "required": ["service"]
}

This agent's own recent access history for one granted service — event type, timestamp, and outcome only. No IP address or other PII beyond what the agent's own accesses already generated, and no credential material.

This is paginated — a page may return fewer than `limit` matches

Each call reads a single page of at most 100 Firestore activity rows for the agent and filters that page down to the requested service — a call can return fewer than limit matching rows (even zero) while still having more history further back, if a noisy OTHER granted service crowds this one out of the current page. Check nextCursor in the response and pass it back as cursor to keep paging; nextCursor is absent once there's nothing left. This bounds the Firestore cost of a single call regardless of how much unrelated activity the agent has generated. credential_history also has its own tighter rate-limit ceiling (20/minute per agent) layered on top of the shared floor every MCP tool call already counts against — see Rate limiting below.

tools/call result — structuredContent
{
  "history": [
    { "eventType": "AGENT_CREDENTIAL_ACCESS", "timestamp": "2026-09-25T10:00:00+00:00", "outcome": "success" }
  ],
  "nextCursor": "eyJ0cyI6ICIyMDI2LTA5LTI1VDEwOjAwOjAwKzAwOjAwIiwgImlkIjogImRvYy0xMjMifQ=="
}

Rate limiting

initialize, tools/list, and notifications/* are protocol bookkeeping, not grant traffic — they skip the per-agent floor limit, the account aggregate, and penalty-box strikes, and count instead against a much more generous mcp_handshake scope (120/minute per agent). They still count against the account-wide aggregate ceiling (with no strike) — credentialed traffic can't bypass the account cap just by only ever calling exempt methods.

Every other method — in particular tools/call — shares the same per-agent floor counter and penalty-box lockout as GET /api/agents/credentials; see that page's Rate limiting section for the full penalty box → floor → aggregate pipeline, and the no-grant miss limit that applies to get_credential exactly as it does to the REST endpoint.

credential_history additionally has its own, tighter ceiling — credential_history (20/minute per agent, env-overridable) — checked before any Firestore read, on top of the shared floor above. Each call can page an expensive activity-log query, so it doesn't get to spend the same generous allowance a cheap single-document read like show_credential gets.

A JSON-RPC batch (a JSON array of request objects) is rate-limited as a single unit — if every message in the batch is handshake-exempt, the WHOLE batch counts against mcp_handshake; if even one message is a real tool call, the whole batch counts against the floor. This transport doesn't implement multi-response batch dispatch, though: a batch body gets a single -32600 error instead of a per-message response —

A JSON-RPC array body
{
  "jsonrpc": "2.0",
  "id": null,
  "error": { "code": -32600, "message": "Batch requests are not supported" }
}

Rate and penalty denials (JSON-RPC error shape)

A floor/penalty-box/handshake/aggregate denial — checked by the shared gate BEFORE your request ever reaches a tool — comes back as a JSON-RPC error (not a tool result) carrying your request's own id, HTTP 429, and a Retry-After header:

429 — floor limit exceeded
{
  "jsonrpc": "2.0",
  "id": 7,
  "error": {
    "code": -32029,
    "message": "Rate limit exceeded.",
    "data": { "code": "RATE_LIMITED", "retryAfter": 12 }
  }
}

-32029 is a Token Vault–specific code in JSON-RPC's reserved "server error" band (-32000…-32099), not one of the JSON-RPC 2.0 base codes. On this endpoint data.code is always RATE_LIMITED — a policy denial (including an ABAC rate_limit rule) is checked INSIDE the tool handler instead, so it's a tool result, not a JSON-RPC error; see Tool errors and structuredContent below. POLICY_DENIED as a -32029 data.code value is specific to /api/proxy/mcp, which has no tools of its own to attach a result to.

Tool errors and structuredContent

A tools/call denial — including an ABAC policy denial, whether or not the denying rule is rate_limit — comes back as an ordinary JSON-RPC result (HTTP 200) with isError: true and a machine-readable structuredContent, so tvault explain <code> (or your own client) can branch on code instead of parsing prose:

200 OK — isError tool result
{
  "jsonrpc": "2.0",
  "id": 7,
  "result": {
    "content": [{ "type": "text", "text": "Error [NO_GRANT]: No grant for service 'stripe'." }],
    "isError": true,
    "structuredContent": { "code": "NO_GRANT", "message": "No grant for service 'stripe'." }
  }
}
codeMeaning
NO_GRANTNo grant for the requested service — deliberately identical whether the grant is missing or the service doesn't exist at all; this never confirms or denies a credential's existence
GRANT_EXPIREDThe grant existed but has lapsed (deleted on this call)
POLICY_DENIEDAn ABAC rule blocked the call — structuredContent.retryAfter is present only for a rate_limit rule
RATE_LIMITEDOver the no-grant miss limit (10/minute) — rate-limited only, never a penalty-box lockout; see Rate limiting
VAULT_LOCKEDThe vault owner has locked the vault
WEBHOOK_NOT_CONFIGUREDNo webhook is bound to this vault
INVALID_ARGUMENTget_credential, show_credential, or credential_history called without service (or with a non-string//-containing one)

structuredContent.retryAfter is present only for RATE_LIMITED and a rate_limit-flavoured POLICY_DENIED.

Policies and revocation

ABAC policies attach to the agent identity, not the transport — time windows, IP allowlists, rate limits, and usage caps all apply to MCP tool calls exactly as to REST. Suspending the agent kills both paths; for OAuth sessions, every token refresh re-checks agent status, and POST /api/mcp-oauth/revoke kills a session immediately.

Ready to try it?

Sign up free with Google — your credentials stay on your own webhook, and the quickstart gets an agent fetching its first credential in about ten minutes.

On this page