Files
sure/docs/hosting/mcp.md
T
efb7cc3935 Tooling for the wealth + tax agent harness (#2848)
* Expose the Statement Vault to external agents over MCP

A user wants to manage patrimonial history — a document-backed record of a
family's wealth where every figure traces back to the statement it came from —
by pointing an external agent harness at Sure. That model belongs in the
harness, not in Sure: it needs numbered build deltas, golden tests and closed
periods that a mutable Postgres row cannot provide.

What Sure was missing was the seam. The Statement Vault already does most of
the work — original bytes retained, SHA-256 dedup, period detection, account
matching with a confidence score, reconciliation against ledger balances, and a
month-by-month coverage map — but it is reachable only from the web UI. An
agent could not archive a document, cite one, or check for gaps.

Adds five preview MCP tools over what already exists, plus a citation grammar
for values the agent writes:

- upload_account_statement, list_account_statements, get_account_statement,
  get_statement_coverage
- record_valuation, whose source citation is parsed rather than trusted:
  ["estimated: "] citation [" (grade: A|B|C)"]. An uncited or free-styled
  value is rejected at the write boundary instead of landing in the ledger
  looking authoritative.

link and reject are deliberately not exposed. Attaching a statement to an
account is the human's decision, and the vault UI is where it is made; the
agent reports the suggested match and stops there.

Assistant.function_classes now takes a user so preview tools stay out of the
default surface. They are hidden from tools/list and not callable by name
without the preference enabled, and the vault tools re-check the manager role
and per-account permissions, since MCP calls never pass through a controller.

Docs: the blueprint this implements, and a guide covering which side owns which
layer, the vocabulary map between the two, the monthly runbook, and the gaps
(non-user holders, non-statement documents, one value per date).

No migrations, no API endpoints, no UI.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JFDp9HhXDeswadu4cxFojn

* Address review feedback on the vault MCP tools

Two non-blocking items from the review pass:

Document why get_statement_coverage reads through accessible_by rather than
writable_by. It reports which documents exist and writes nothing, so read
access is the right bar — and tightening it would hide coverage gaps from
people who can already see the figures those gaps sit behind. The comment
exists so a future refactor doesn't "fix" it.

Close the acknowledged verification gap with tests rather than a one-off
manual check. The review noted that nothing proved a real vault payload
serializes cleanly out through tools/call — vault responses are richer than
the other tools' output, with nested account hashes, decimal balances, dates
and a compacted hash. Two integration tests now drive the real /mcp endpoint
end to end against a real AccountStatement: one listing it, one uploading
bytes and reading back the SHA-256. Permanent regression coverage instead of
a smoke test someone has to remember to repeat.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JFDp9HhXDeswadu4cxFojn

* docs(llm-guides): replace patrimonial blueprint with its final revision

Swap the embedded early draft for the authoritative final revision of the
wealth + tax modelling blueprint (MIT © 2026 diegomarino):

- rename the domain vocabulary: patrimonial -> wealth, fiscal -> tax
  (tax_data/, the tax layer, tax_runner)
- add §9.5 (the intel file: shape, generation, and the capture loop)
- tighten worked examples down to placeholders
- add the MIT header; keep the in-repo NOTE block (adapted to the new
  vocabulary) and the filename untouched so cross-links don't break

* docs(llm-guides): align agent-harness guide with blueprint + fix reconcile semantics

Follow the blueprint rename (patrimonial -> wealth, fiscal -> tax,
fiscal_data/ -> tax_data/, "Phase 7 (fiscal layer)" -> "(the tax layer)")
so the two docs stop disagreeing on vocabulary.

Correct the reconciliation mapping, which conflated two different invariants:

- blueprint reconcile-or-abort (§7 pass 3) is parse-integrity (parsed parts
  == the document's own printed total); Sure's reconciliation_checks is
  ledger agreement (statement balances vs the ledger). Sure has no
  parse-integrity check and never aborts.
- opening_balance / closing_balance are user-entered, not auto-extracted, so
  over MCP reconciliation is "unavailable" until a human fills them.
- tolerance differs: blueprint 1.00/account-period vs Sure's fixed 0.01.

State in the ownership table, the invariants section, the vocabulary map and
the monthly runbook that parse-integrity and the abort belong to the harness
extractor.

* Correct the vault tools' reconciliation claims and citation parsing

Review findings from @diegomarino, all verified against the code before
changing anything.

The reconciliation claim was the serious one. get_account_statement told
agents the checks were "the trustworthy part" and returned "the balances read
off it" — but nothing reads balances off a document. MetadataDetector never
touches them and create_from_prepared_upload! never sets them; they are
user-editable fields in the Statement Vault UI. So a statement archived over
MCP always came back with an empty check list, which an agent could easily
read as "the document agrees with the ledger" when it means "nobody has
entered the figures". The description now says so, and the payload carries a
reconciliation_note spelling it out for anything reading only the JSON. Also
noted that these checks are ledger agreement, not parse integrity: nothing
here verifies a document's parts sum to its printed total.

Provenance::Citation had two patterns disagreeing about spacing. GRADE_SUFFIX
allowed "(grade:A)" but FORMAT required exactly one space, so that citation
passed the pre-check and then parsed as ungraded with the grade swallowed into
the text — silently discarding the reliability the caller supplied, which is
the one thing this parser exists to prevent.

list_account_statements downcases content_sha256 before querying. The column
is constrained to lowercase hex, so uppercase input could never match, and an
agent would read the empty result as "not archived" and upload a duplicate.
Its period filters are renamed overlapping_from / overlapping_until, since
they match on overlap and the old names claimed otherwise to anyone reading
the schema without the descriptions. has_more now explains that there is no
cursor and the way forward is a bigger limit or narrower filters.

record_valuation no longer overwrites the entry's notes. Re-recording a date
would destroy a note a person had written there. Nothing is removed now: an
identical citation is a no-op, a changed one is appended, and the trail of
what was cited when survives. Detecting "did this tool write that line?" is
not possible — almost any prose parses as a valid ungraded citation — so the
code does not guess.

Minor: accept urlsafe base64 on upload, and explain in the code why
record_valuation checks the account ACL rather than the vault manager role, so
nobody "tightens" it into the wrong permission later.

Tests cover each: the grade-spacing cases both ways, uppercase SHA lookup,
overlap window boundaries, note preservation and no-stacking, the unavailable
reconciliation note appearing and disappearing, and — per the review — that
the download URL's signed id actually expires, rather than trusting the
description's claim.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JFDp9HhXDeswadu4cxFojn

* Repair a bad merge in the MCP controller test

The merge of main spliced the incoming `tools/call executes
update_transaction` test into the middle of the upload round-trip test,
before its closing `end`. That left the file one `end` short, so it did
not parse — taking out both `ci / lint` (Lint/Syntax) and `ci / test_unit`
(the whole file failed to load).

Restores the missing `end`. Both tests are kept as their authors wrote
them; nothing else changes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JFDp9HhXDeswadu4cxFojn

* docs: use wealth history wording (#2885)

* Stop the vault tools promising verification they don't perform

Three findings from the automated review passes, all confirmed against the
code before changing anything.

The download URL was dead on arrival for the caller it was built for. Sure
serves stored files through Active Storage controllers that
config/initializers/active_storage_authorization.rb gates on
`viewable_by?(Current.user)` — a signed-in browser session. An MCP client has
a bearer token and no session, so following the URL would have redirected to
sign-in. Removed it rather than leaving a link that cannot work, and the
description now points at search_family_files or the vault UI.

Coverage called a month `covered` when a document merely existed. An
unreconciled statement is not mismatched, so it took the `covered` branch, and
the payload carried nothing to correct the reading — the same "advertised
verification that never happened" bug fixed last round in
get_account_statement, in a second place. Months now carry their own
reconciliation_status, and the description says covered means presence, not
agreement.

Listing filtered visibility after limiting. Beyond underfilling a page, with
no cursor and a 100-row cap an accessible statement behind enough newer
invisible ones was unreachable. Visibility now lives in the query, mirroring
viewable_by? for a statement manager.

Also: rescue unexpected upload failures into a tool error instead of a raw
exception string, derive the documented size limit from MAX_FILE_SIZE, list
every coverage status in mcp.md, and cover the failed-reconciliation and
base64-normalisation branches.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JFDp9HhXDeswadu4cxFojn

* Keep storage exception detail out of the MCP response

The upload_failed message interpolated the exception text, which crosses out
to an external agent. A storage failure can carry bucket names, object keys,
paths or request details, so the agent now gets a fixed message and the
exception stays in the server log. The test asserts the absence of detail
rather than pinning the leaked string into the contract.

Also fixes a test that did not test what it claimed: the urlsafe-base64 case
used a fixture encoding to plain base64, so it exercised the padding branch
and never the "-_" translation. It now uses content whose encoding contains
both characters and asserts that up front.

Renames "rejects content that decodes to zero bytes" to "rejects blank
content", which is what it actually covers — Base64.strict_encode64("") is
"", which is blank and returns before the decoder runs, so invalid_content is
correct and empty_file is not reachable from this path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JFDp9HhXDeswadu4cxFojn

* docs: align wealth blueprint review feedback

* Correct the harness runbook: parse before publishing

The guide told an implementer to archive each document to Sure first and
work from there. That strands them. Sure never returns a document's bytes
over MCP — Active Storage serves stored files only to a signed-in browser
session — and there is no text fallback either, because statements archived
through upload_account_statement never enter the vector store, so
search_family_files cannot see them. A statement in Sure is metadata to an
agent and nothing more.

That blocks exactly three blueprint steps, all of them operating on bank and
broker statements: the extractors, the parts-vs-printed-total check, and the
glyph decoder. Everything else it parses — tax returns, capital accounts,
annual accounts — the harness already holds locally.

So the order inverts: the harness ingests into its own vault, extracts there
with the whole file in reach, and publishes to Sure afterwards. This restores
principle 8 rather than bending it — the recurring pipeline reads from the
canonical store, and treating Sure as canonical forced a re-fetch the
architecture never sanctioned. Both sides hash the same bytes, so the SHA-256
verifies Sure holds the identical document without moving it.

Writes down the two consequences: a statement uploaded straight into Sure's
UI can be known but never parsed (reliability C or PENDING until a copy
reaches the harness), and neither vault backs up the other.

Also drops a stale tools-table row still advertising the 15-minute download
URL removed earlier, corrects get_account_statement's description where it
suggested search_family_files as a fallback it cannot be, and disambiguates
"the vault" in the MCP tool table, which is what misled me in the first place.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JFDp9HhXDeswadu4cxFojn

---------

Signed-off-by: Juan José Mata <juanjo.mata@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: diegomarino <diegomarino@users.noreply.github.com>
Co-authored-by: Sure Admin (bot) <sure-admin@splashblot.com>
2026-08-04 23:33:01 +02:00

12 KiB

MCP Server for External AI Assistants

Sure includes a Model Context Protocol (MCP) server endpoint that allows external AI assistants like Claude Desktop, GPT agents, or custom AI clients to query your financial data.

What is MCP?

Model Context Protocol is a JSON-RPC 2.0 protocol that enables AI assistants to access structured data and tools from external applications. Instead of copying and pasting financial data into a chat window, your AI assistant can directly query Sure's data through a secure API.

This is useful when:

  • You want to use an external AI assistant (Claude, GPT, custom agents) to analyze your Sure financial data
  • You prefer to keep your LLM provider separate from Sure
  • You're building custom AI agents that need access to financial tools

Prerequisites

For legacy bearer-token authentication, set these two environment variables:

Variable Description Example
MCP_API_TOKEN Bearer token for authentication your-secret-token-here
MCP_USER_EMAIL Email of the Sure user whose data the assistant can access user@example.com

Both variables are required for the legacy token flow. OAuth clients using the MCP discovery and dynamic registration endpoints do not need these variables.

Generating a secure token

Generate a random token for MCP_API_TOKEN:

# macOS/Linux
openssl rand -base64 32

# Or use any secure password generator

Choosing the user

The MCP_USER_EMAIL must match an existing Sure user's email address. The AI assistant will have access to all financial data for that user's family.

Caution

The AI assistant will have read access to all financial data for the specified user. Only set this for users you trust with your AI provider.

Configuration

Docker Compose

Add the environment variables to your compose.yml:

x-rails-env: &rails_env
  MCP_API_TOKEN: your-secret-token-here
  MCP_USER_EMAIL: user@example.com

Both web and worker services inherit this configuration.

Kubernetes (Helm)

Add the variables to your values.yaml or set them via Secrets:

env:
  MCP_API_TOKEN: your-secret-token-here
  MCP_USER_EMAIL: user@example.com

Or create a Secret and reference it:

envFrom:
  - secretRef:
      name: sure-mcp-credentials

Protocol Details

The MCP endpoint is available at:

POST /mcp

Authentication

MCP supports OAuth authorization-code flow for clients such as Claude Code. Clients should discover the protected-resource metadata, register dynamically, request the advertised read_write scope, and send the resulting access token as a Bearer token. Dynamically registered clients are assigned this scope so their tokens can authenticate to MCP.

For self-hosted deployments or clients without OAuth support, requests may use the legacy MCP_API_TOKEN as a Bearer token:

Authorization: Bearer <MCP_API_TOKEN>

Supported Methods

Sure implements the following JSON-RPC 2.0 methods:

Method Description
initialize Protocol handshake, returns server info and capabilities
tools/list Lists available financial tools with schemas
tools/call Executes a tool with provided arguments

Available Tools

The MCP endpoint exposes these financial tools:

Tool Description
get_transactions Retrieve transaction history with filtering
get_accounts Get account information and balances
get_holdings Query investment holdings
get_balance_sheet Current financial position (assets, liabilities, net worth)
get_income_statement Income and expenses over a period
import_bank_statement Import bank statement data
search_family_files Search documents uploaded through the import flow. Note this is the vector-store document index, not the Statement Vault — statements archived via upload_account_statement are not searchable through it

These are the same tools used by Sure's builtin AI assistant.

Preview Tools

These additional tools appear only when the MCP user has opted into preview features (Settings → Preferences). Until then they are absent from tools/list, and calling one by name returns an "Unknown tool" error. The Statement Vault tools additionally require the user to be an admin or member, matching the permissions enforced in the web UI.

Tool Description
upload_account_statement Store a statement document (PDF/CSV/XLSX) in the Statement Vault; deduplicates by SHA-256
list_account_statements List vault documents with their SHA-256, period, linked account and review status
get_account_statement One statement's details and its reconciliation checks against the ledger — present only once someone has entered the statement's opening/closing balances in the web UI, since nothing extracts them from the document. Does not return the file: stored documents are served only to a signed-in browser session
get_statement_coverage Month-by-month statement coverage for an account: covered, missing, mismatched, ambiguous, duplicate, not_expected, each with a reconciliation status
record_valuation Record an account's value on a date, with a required source citation

They exist for agents that maintain a document-backed record of a family's wealth over time. See Wealth history with an external agent harness.

Example Requests

Initialize

Handshake to verify protocol version and capabilities:

curl -X POST https://your-sure-instance/mcp \
  -H "Authorization: Bearer your-secret-token" \
  -H "Content-Type: application/json" \
  -d '{
    "jsonrpc": "2.0",
    "id": 1,
    "method": "initialize"
  }'

Response:

{
  "jsonrpc": "2.0",
  "id": 1,
  "result": {
    "protocolVersion": "2025-03-26",
    "capabilities": {
      "tools": {}
    },
    "serverInfo": {
      "name": "sure",
      "version": "1.0"
    }
  }
}

List Tools

Get available tools with their schemas:

curl -X POST https://your-sure-instance/mcp \
  -H "Authorization: Bearer your-secret-token" \
  -H "Content-Type: application/json" \
  -d '{
    "jsonrpc": "2.0",
    "id": 2,
    "method": "tools/list"
  }'

Response includes tool names, descriptions, and JSON schemas for parameters.

Call a Tool

Execute a tool to get transactions:

curl -X POST https://your-sure-instance/mcp \
  -H "Authorization: Bearer your-secret-token" \
  -H "Content-Type: application/json" \
  -d '{
    "jsonrpc": "2.0",
    "id": 3,
    "method": "tools/call",
    "params": {
      "name": "get_transactions",
      "arguments": {
        "start_date": "2024-01-01",
        "end_date": "2024-01-31"
      }
    }
  }'

Response:

{
  "jsonrpc": "2.0",
  "id": 3,
  "result": {
    "content": [
      {
        "type": "text",
        "text": "[{\"id\":\"...\",\"amount\":-45.99,\"date\":\"2024-01-15\",\"name\":\"Coffee Shop\"}]"
      }
    ]
  }
}

Security Considerations

Transient Session Isolation

The MCP controller creates a transient session for each request. This prevents session state leaks that could expose other users' data if the Sure instance is using impersonation features.

Each MCP request:

  1. Authenticates the token
  2. Loads the user specified in MCP_USER_EMAIL
  3. Creates a temporary session scoped to that user
  4. Executes the tool call
  5. Discards the session

This ensures the AI assistant can only access data for the intended user.

Pipelock Security Scanning

For production deployments, we recommend using Pipelock to scan MCP traffic for security threats.

Pipelock provides:

  • DLP scanning: Detects secrets being exfiltrated through tool calls
  • Prompt injection detection: Identifies attempts to manipulate the AI
  • Tool poisoning detection: Prevents malicious tool call sequences
  • Policy enforcement: Block or warn on suspicious patterns
  • Signed receipts: Produces verifiable evidence for mediated MCP decisions when the flight recorder is configured with storage and a signing key

See the Pipelock documentation and the example configuration in compose.example.ai.yml for setup instructions.

Network Security

The /mcp endpoint is exposed on the same port as the web UI (default 3000). For hardened deployments:

Docker Compose:

  • The MCP endpoint is protected by the MCP_API_TOKEN but is reachable on port 3000
  • For additional security, use Pipelock's MCP reverse proxy (port 8889) which adds scanning
  • See compose.example.ai.yml for a Pipelock configuration

Kubernetes:

  • Use NetworkPolicies to restrict access to the MCP endpoint
  • Route external agents through Pipelock's MCP reverse proxy
  • See the Helm chart documentation for Pipelock ingress setup

Production Deployment

For a production-ready setup with security scanning:

  1. Download the example configuration:

    curl -o compose.ai.yml https://raw.githubusercontent.com/we-promise/sure/main/compose.example.ai.yml
    curl -o pipelock.example.yaml https://raw.githubusercontent.com/we-promise/sure/main/pipelock.example.yaml
    
  2. Set your MCP credentials in .env:

    MCP_API_TOKEN=your-secret-token
    MCP_USER_EMAIL=user@example.com
    
  3. Start the stack:

    docker compose -f compose.ai.yml up -d
    
  4. Connect your AI assistant to the Pipelock MCP proxy:

    http://your-server:8889
    

The Pipelock proxy (port 8889) scans all MCP traffic before forwarding to Sure's /mcp endpoint.

Connecting AI Assistants

Claude Desktop

Configure Claude Desktop to use Sure's MCP server:

  1. Open Claude Desktop settings
  2. Add a new MCP server
  3. Set the endpoint to http://your-server:8889 (if using Pipelock) or http://your-server:3000/mcp
  4. Add the authorization header: Authorization: Bearer your-secret-token

Custom Agents

Any AI agent that supports JSON-RPC 2.0 can connect to the MCP endpoint. The agent should:

  1. Send a POST request to /mcp
  2. Include the Authorization: Bearer <token> header
  3. Use the JSON-RPC 2.0 format for requests
  4. Handle the protocol methods: initialize, tools/list, tools/call

Troubleshooting

"MCP endpoint not configured" error

Symptom: Requests return HTTP 503 with "MCP endpoint not configured"

Fix: Ensure both MCP_API_TOKEN and MCP_USER_EMAIL are set as environment variables and restart Sure.

"unauthorized" error

Symptom: Requests return HTTP 401 with "unauthorized"

Fix: Verify the Authorization header contains the correct token: Bearer <MCP_API_TOKEN>

"MCP user not configured" error

Symptom: Requests return HTTP 503 with "MCP user not configured"

Fix: The MCP_USER_EMAIL does not match an existing user. Check that:

  • The email is correct
  • The user exists in the database
  • There are no typos or extra spaces

Pipelock connection refused

Symptom: AI assistant cannot connect to Pipelock's MCP proxy (port 8889)

Fix:

  1. Verify Pipelock is running: docker compose ps pipelock
  2. Check Pipelock health: docker compose exec pipelock /pipelock healthcheck --addr 127.0.0.1:8888
  3. Verify the port is exposed in your compose.yml

See Also