MCP Endpoint
Reference for POST /api/agents/mcp — the MCP server exposing list_credentials and get_credential tools, with tvagent_ key or OAuth 2.1 session auth.
POST /api/agents/mcp
Token Vault's own MCP server. Any MCP client (Claude Code, Claude.ai, ChatGPT, Cursor) can connect over streamable HTTP and use standard tool calls to work with the agent's granted credentials.
claude mcp add --transport http token-vault https://api.tokenvault.uk/api/agents/mcpAuthentication
Two interchangeable identities — both resolve to the same agent, grants, and policies:
- OAuth 2.1 (recommended for MCP clients) — the client discovers Token Vault's authorization server from a
401, registers itself, and obtains atvsess_…session token in-client. No key is pasted anywhere. See OAuth 2.1 for MCP Clients. - Static key —
Authorization: Bearer tvagent_…for server-side agents you configure by hand.
Auth is required on every JSON-RPC method, including initialize. An unauthenticated (or
invalid) call to initialize gets the same 401 + WWW-Authenticate discovery challenge as any
other method:
WWW-Authenticate: Bearer resource_metadata="https://api.tokenvault.uk/.well-known/oauth-protected-resource"In practice this changes nothing for a real client — the same header you already send on tools/call
is just checked one method earlier. It exists because this endpoint shares its rate-limit/penalty-box
state with the REST credentials endpoint, which needs a resolved
identity to key on before any of the layers below can run.
Tools
| Tool | Arguments | Returns |
|---|---|---|
list_credentials | — | The agent's active grants, enriched with service metadata (description, display name, tags, token type, expiry, whether it carries a refresh token) — no credential material |
get_credential | service | A signed, short-lived credential URL pointing at your webhook's /v1/credential |
whoami | — | This agent's own identity: id, name, status, owner, active grants, and the names of any ABAC policies attached to it — no credential material |
show_credential | service | Metadata for ONE granted service — the same fields a list_credentials entry carries — without retrieving the credential itself |
credential_history | service, limit (optional, max 50), cursor (optional) | This agent's own recent access history for one granted service: event type, timestamp, and outcome — no credential material |
Why a URL and not the credential?
The MCP transport runs through Token Vault, and Token Vault never carries credential bytes. get_credential therefore returns a one-time signed URL; the agent fetches it and receives the credential directly from your webhook — the same custody path as the REST endpoint's 307 redirect. Every other tool on this page — list_credentials, whoami, show_credential, credential_history — only ever returns metadata Token Vault already holds in its own control plane (grants, agent identity, activity log). None of it is the credential, derived from the credential, or a substitute for fetching it.
list_credentials
No arguments. Each granted service is enriched with Firestore-owned metadata (description, displayName, tags) and a whitelisted slice of your webhook's own list response (tokenType, expiryTime, hasRefreshToken). The webhook lookup is best-effort — if it fails, the result silently falls back to the metadata-only fields instead of erroring.
{
"grants": [
{
"serviceName": "github",
"grantExpiresAt": null,
"refreshPolicy": "none",
"description": "GitHub PAT for CI",
"displayName": "GitHub",
"tags": ["ci"],
"tokenType": "pat",
"expiryTime": null,
"hasRefreshToken": false
}
]
}whoami
{ "type": "object", "properties": {}, "required": [] }Returns this agent's own id, name, status, owner, active grants, and attached policy names — never the owner's email, and never a policy's rule configuration, only its name and rule type(s).
{
"id": "agent-1",
"name": "CI Bot",
"status": "active",
"ownerId": "firebase-uid",
"grants": [
{ "serviceName": "github", "grantExpiresAt": null, "refreshPolicy": "none" }
],
"policies": [
{ "name": "Business Hours", "ruleTypes": ["time_window"] }
]
}show_credential
{
"type": "object",
"properties": {
"service": { "type": "string", "description": "The granted service name to show metadata for." }
},
"required": ["service"]
}The same fields a list_credentials entry carries, for one service — metadata only, never the credential. An ungranted service returns the identical, deliberately ambiguous NO_GRANT error get_credential uses, so this can't be used to enumerate services the agent was never granted.
{
"serviceName": "github",
"grantExpiresAt": null,
"refreshPolicy": "none",
"displayName": "GitHub",
"tokenType": "pat",
"hasRefreshToken": false
}credential_history
{
"type": "object",
"properties": {
"service": { "type": "string", "description": "The granted service name to show access history for." },
"limit": {
"type": "integer",
"description": "Maximum number of history rows to return per page (default 20, max 50).",
"minimum": 1,
"maximum": 50
},
"cursor": {
"type": "string",
"description": "Opaque pagination cursor from a previous call's nextCursor, to fetch the next page."
}
},
"required": ["service"]
}This agent's own recent access history for one granted service — event type, timestamp, and outcome only. No IP address or other PII beyond what the agent's own accesses already generated, and no credential material.
This is paginated — a page may return fewer than `limit` matches
Each call reads a single page of at most 100 Firestore activity rows for the agent and filters that
page down to the requested service — a call can return fewer than limit matching rows (even zero)
while still having more history further back, if a noisy OTHER granted service crowds this one out of
the current page. Check nextCursor in the response and pass it back as cursor to keep paging;
nextCursor is absent once there's nothing left. This bounds the Firestore cost of a single call
regardless of how much unrelated activity the agent has generated. credential_history also has its
own tighter rate-limit ceiling (20/minute per agent) layered on top of the shared floor every MCP tool
call already counts against — see Rate limiting below.
{
"history": [
{ "eventType": "AGENT_CREDENTIAL_ACCESS", "timestamp": "2026-09-25T10:00:00+00:00", "outcome": "success" }
],
"nextCursor": "eyJ0cyI6ICIyMDI2LTA5LTI1VDEwOjAwOjAwKzAwOjAwIiwgImlkIjogImRvYy0xMjMifQ=="
}Rate limiting
initialize, tools/list, and notifications/* are protocol bookkeeping, not grant traffic — they
skip the per-agent floor limit, the account aggregate, and penalty-box strikes, and count instead
against a much more generous mcp_handshake scope (120/minute per agent). They still count
against the account-wide aggregate ceiling (with no strike) — credentialed traffic can't bypass
the account cap just by only ever calling exempt methods.
Every other method — in particular tools/call — shares the same per-agent floor counter and
penalty-box lockout as GET /api/agents/credentials; see that page's Rate limiting
section for the full penalty box → floor → aggregate pipeline, and the no-grant miss limit that
applies to get_credential exactly as it does to the REST endpoint.
credential_history additionally has its own, tighter ceiling — credential_history (20/minute
per agent, env-overridable) — checked before any Firestore read, on top of the shared floor above.
Each call can page an expensive activity-log query, so it doesn't get to spend the same generous
allowance a cheap single-document read like show_credential gets.
A JSON-RPC batch (a JSON array of request objects) is rate-limited as a single unit — if every
message in the batch is handshake-exempt, the WHOLE batch counts against mcp_handshake; if even one
message is a real tool call, the whole batch counts against the floor. This transport doesn't
implement multi-response batch dispatch, though: a batch body gets a single -32600 error instead of
a per-message response —
{
"jsonrpc": "2.0",
"id": null,
"error": { "code": -32600, "message": "Batch requests are not supported" }
}Rate and penalty denials (JSON-RPC error shape)
A floor/penalty-box/handshake/aggregate denial — checked by the shared gate BEFORE your request
ever reaches a tool — comes back as a JSON-RPC error (not a tool result) carrying your request's
own id, HTTP 429, and a Retry-After header:
{
"jsonrpc": "2.0",
"id": 7,
"error": {
"code": -32029,
"message": "Rate limit exceeded.",
"data": { "code": "RATE_LIMITED", "retryAfter": 12 }
}
}-32029 is a Token Vault–specific code in JSON-RPC's reserved "server error" band
(-32000…-32099), not one of the JSON-RPC 2.0 base codes. On this endpoint data.code is
always RATE_LIMITED — a policy denial (including an ABAC rate_limit rule) is checked INSIDE the
tool handler instead, so it's a tool result, not a JSON-RPC error; see
Tool errors and structuredContent below. POLICY_DENIED as a
-32029 data.code value is specific to
/api/proxy/mcp, which has
no tools of its own to attach a result to.
Tool errors and structuredContent
A tools/call denial — including an ABAC policy denial, whether or not the denying rule is
rate_limit — comes back as an ordinary JSON-RPC result (HTTP 200) with isError: true and a
machine-readable structuredContent, so tvault explain <code> (or your own client) can branch on
code instead of parsing prose:
{
"jsonrpc": "2.0",
"id": 7,
"result": {
"content": [{ "type": "text", "text": "Error [NO_GRANT]: No grant for service 'stripe'." }],
"isError": true,
"structuredContent": { "code": "NO_GRANT", "message": "No grant for service 'stripe'." }
}
}code | Meaning |
|---|---|
NO_GRANT | No grant for the requested service — deliberately identical whether the grant is missing or the service doesn't exist at all; this never confirms or denies a credential's existence |
GRANT_EXPIRED | The grant existed but has lapsed (deleted on this call) |
POLICY_DENIED | An ABAC rule blocked the call — structuredContent.retryAfter is present only for a rate_limit rule |
RATE_LIMITED | Over the no-grant miss limit (10/minute) — rate-limited only, never a penalty-box lockout; see Rate limiting |
VAULT_LOCKED | The vault owner has locked the vault |
WEBHOOK_NOT_CONFIGURED | No webhook is bound to this vault |
INVALID_ARGUMENT | get_credential, show_credential, or credential_history called without service (or with a non-string//-containing one) |
structuredContent.retryAfter is present only for RATE_LIMITED and a rate_limit-flavoured
POLICY_DENIED.
Policies and revocation
ABAC policies attach to the agent identity, not the transport — time windows, IP allowlists, rate limits, and usage caps all apply to MCP tool calls exactly as to REST. Suspending the agent kills both paths; for OAuth sessions, every token refresh re-checks agent status, and POST /api/mcp-oauth/revoke kills a session immediately.
Ready to try it?
Sign up free with Google — your credentials stay on your own webhook, and the quickstart gets an agent fetching its first credential in about ten minutes.
Agent Credentials API
Reference for GET /api/agents/credentials — auth forms, the 307 redirect contract, list mode, response schemas, and the full error catalogue.
MCP Proxy Endpoint
Reference for POST /api/proxy/mcp — proxy key auth forms, webhook credential injection, and the tv_session upstream auth type.