mcp-proof / evaluation
Effect-aware conformance
Can declared MCP effects be verified against independently observed external state?
Declared MCP semantics can be misleading even when tool responses and tool names look correct. Each experiment compares a declared or inferred property against ground truth established outside the audited MCP surface — an out-of-band state observer and a probe that exercises created objects. E1 is the detection experiment; E2 validates that the persistence/authority distinction is operationally real (construct validation); E3 measures effectiveness after lifecycle events (residual authority); the case studies take the instrument to servers we did not build.
E1–E3 are controlled-testbed measurements of detection behaviour on planted inconsistencies — not production prevalence. Ground truth is read out-of-band from the state store, never through the audited MCP surface. The case studies run the same instrument against published third-party servers.
Artifacts: effect-report-honest.html effect-report-silent-keymint.html summary.md e1.json e2.json e3.json case-memory.json case-filesystem.json case-sqlite.json
The detection experiment: can effect observation catch annotation lies that a response-only auditor and a name heuristic miss?
Only effect observation detects all three response-invisible lies. Each baseline reaches recall 0.333 with a structural blind spot the observation channel cannot fix.
| Detector | Precision | Recall | F1 |
|---|---|---|---|
| Effect-awareout-of-band state observation · TP 6 · FP 0 · FN 0 | 1.000 | 1.000 | 1.000 |
| Response-leveltool responses only (current auditor) · TP 2 · FP 0 · FN 4 | 1.000 | 0.333 | 0.500 |
| Name heuristictool names/descriptions only · TP 2 · FP 0 · FN 4 | 1.000 | 0.333 | 0.500 |
aggregated over the honest server + six single-lie variants; per-lie detections below
| Planted lie | Lying tool | Effect-aware | Response-level | Name heuristic |
|---|---|---|---|---|
| effect visible in the response | ||||
lie-readonly | save_note | caught | caught | missed |
hide-destructive | delete_note | caught | caught | caught |
webhook-readonly | create_webhook | caught | missed | caught |
| effect invisible in the response — only observation can see it | ||||
silent-keymint | get_note | caught | missed | missed |
shadow-webhook | list_notes | caught | missed | missed |
phantom-write | ping | caught | missed | missed |
● caught○ missed — highlighted rows are the three response-invisible lies
Inspect evidence — per-variant detections
| Server variant | Ground truth (lying tool) | Effect-aware flagged | Response-level flagged | Name heuristic flagged |
|---|---|---|---|---|
honest | — | — | — | — |
lie-readonly | save_note | save_note | save_note | — |
hide-destructive | delete_note | delete_note | delete_note | delete_note |
webhook-readonly | create_webhook | create_webhook | — | create_webhook |
silent-keymint | get_note | get_note | — | — |
shadow-webhook | list_notes | list_notes | — | — |
phantom-write | ping | ping | — | — |
One audit per variant; every detector reads the same execution, so differences come from the observation channel, not from run-to-run noise. The honest variant measures false positives — no detector flagged anything on it.
Oracle: the testbed's mutation ledger — each variant plants exactly one annotation lie in one tool. Fairness: all three detectors read the same execution records; they differ only in which channel they may consult (external state, response text, or metadata).
Construct validation, not a benchmark: is the persistence/authority distinction operationally real — does probe-backed classification separate authority-bearing objects from ordinary persistent objects where “persistent ⇒ authority” and a credential-name keyword go wrong?
The probe classifies by use, not appearance — it is correct on the decoy note named api_key_backup (fools both baselines' signals) and on a credential that never persisted.
| Classifier | Accuracy | TP | FP | TN | FN |
|---|---|---|---|---|---|
| Probe (exercise)attempts to use the object | 1.000 | 5 | 0 | 3 | 0 |
| Name keywordcredential-looking names count as authority | 0.875 | 5 | 1 | 2 | 0 |
| Persistence ⇒ authoritypersistent objects count as authority | 0.625 | 4 | 2 | 1 | 1 |
object corpus spans the persistent × authority quadrants; per-object classifications below
| Object | Persistent | Authority (truth) | Persistence says | Name says | Probe says | Note |
|---|---|---|---|---|---|---|
api_keys/key_0001 | true | true | true | true | true | — |
api_keys/key_0002 | true | true | true | true | true | — |
webhooks/wh_0001 | true | true | true | true | true | — |
share_links/share_0001 | true | true | true | true | true | — |
notes/welcome | true | false | true (disagrees with ground truth) | false | false | — |
notes/api_key_backup | true | false | true (disagrees with ground truth) | true (disagrees with ground truth) | false | decoy: credential-looking note name |
notes/scratch | false | false | false | false | false | ephemeral note (created, then deleted in-plan) |
api_keys/key_0003 | false | true | false (disagrees with ground truth) | true | true | counterexample: credential minted then removed out-of-band — authority without persistence |
● true○ false prediction disagrees with ground truth
Oracle: by-construction authority labels — a table's objects are authority-bearing iff the testbed's real authorization rule accepts them as credentials. Probing time: each object is exercised at creation time, while it genuinely exists; persistence is judged against the final snapshot.
A lifecycle measurement: after an event (grant revoked, key revoked, key deleted, TTL expired), does an exercise probe report an object's true effectiveness where “it is still listed” and “its grant is still active” do not — and why authorized_by cannot be read as depends_on?
In grant_revoked the key remains effective although the grant that authorized it is gone — residual authority. The delegation-centric view calls it dead (the one false-ineffective); the probe does not, and existence never distinguishes a live key from a listed-but-dead one.
| Detector | Accuracy | False-effective | False-ineffective |
|---|---|---|---|
| Probe (exercise)attempts to use the object | 1.000 | 0 | 0 |
| Existence (inventory)listed ⇒ effective | 0.500 | 3 | 0 |
| Delegation-centricgrant active ⇒ effective | 0.333 | 3 | 1 — residual authority missed |
false-ineffective is the dangerous error: believing an object is neutralized while it still works
| Lifecycle scenario | Truth | Existence says | Delegation says | Probe says | Note |
|---|---|---|---|---|---|
no_event | effective | effective | effective | effective | key minted under an active grant, nothing revoked |
grant_revoked | effective | effective | dead (disagrees with ground truth) | effective | the authorizing grant is revoked; the key still works (residual authority) residual authority |
key_revoked | dead | effective (disagrees with ground truth) | effective (disagrees with ground truth) | dead | the key itself is revoked; it is still listed but dead |
key_unlisted | dead | dead | effective (disagrees with ground truth) | dead | the key row is deleted; not listed and dead |
ttl_expired | dead | effective (disagrees with ground truth) | effective (disagrees with ground truth) | dead | the key's TTL has passed; still listed but dead |
grant_revoked_cascade | dead | effective (disagrees with ground truth) | dead | dead | grant revoked on a server that cascades; the key is dead |
● effective○ dead prediction disagrees with ground truth
Inspect evidence — per-scenario probe transcript
| Scenario | Key row | Grant | Probe status | Probe basis |
|---|---|---|---|---|
no_event | listed | active | effective | exercised api_keys.secret: server auth rule accepts it (status/expiry checked) |
grant_revoked | listed | revoked | effective | exercised api_keys.secret: server auth rule accepts it (status/expiry checked) |
key_revoked | listed | active | ineffective | exercised api_keys.secret: server auth rule rejects it (status/expiry checked) |
key_unlisted | absent | active | ineffective | object not listed; nothing to exercise |
ttl_expired | listed | active | ineffective | exercised api_keys.secret: server auth rule rejects it (status/expiry checked) |
grant_revoked_cascade | listed | revoked | ineffective | exercised api_keys.secret: server auth rule rejects it (status/expiry/grant checked) |
The probe basis is the evidence line the exercise probe recorded: which credential column it exercised and which rule accepted or rejected it. The key is api_keys/key_0001 in every scenario; each scenario runs on a fresh server.
Oracle: intended effectiveness per scenario, by construction — modelled on documented provider behaviour (a revoked grant leaving a minted key alive; keys listed after revocation). Event application: lifecycle events are applied out-of-band, never through the audited tools.
Does the instrument work on servers we did not build? 3 published MCP servers — official reference servers and a community server — across three store types (JSONL file, directory tree, SQLite) and two annotation profiles (fully annotated, none declared).
No annotation contradiction on any of the 3 audited servers — the honest result. The declarations that could be checked held under observation (readOnly tools wrote nothing; deletes were declared destructive; idempotent claims held under an actual repeat), and where nothing was declared the checks SKIPped instead of inventing a verdict.
| Server (pinned) | Store · observer | Annotations | Observed effects | Effect checks |
|---|---|---|---|---|
@modelcontextprotocol/server-memory@2026.8.31modelcontextprotocol (official reference server) | knowledge-graph JSONL file (MEMORY_FILE_PATH)an out-of-band parse of the JSONL store into per-entity/per-relation objects | 9/9 tools3 readOnly · 3 destructive · 6 idempotent | 2c / 2u / 2dover 10 calls | EFF-01 PASSEFF-02 PASSEFF-03 PASSEFF-06 SKIP effect-evidence report · raw JSON |
@modelcontextprotocol/server-filesystem@2026.8.31modelcontextprotocol (official reference server) | jailed directory treethe stock FilesystemObserver over the jail root | 14/14 tools10 readOnly · 3 destructive · 2 idempotent | 3c / 1u / 0dover 16 calls | EFF-01 PASSEFF-02 PASSEFF-03 PASSEFF-06 SKIP effect-evidence report · raw JSON |
@executeautomation/database-server@1.1.0ExecuteAutomation (community) | SQLite database filethe stock SqliteObserver | none declared0 readOnly · 0 destructive · 0 idempotent | 3c / 1u / 1dover 10 calls | EFF-01 SKIPEFF-02 PASSEFF-03 SKIPEFF-06 SKIP effect-evidence report · raw JSON |
Scope of this validation: these runs exercise the observation half of the lane for real — out-of-band per-object effect attribution and EFF-01/02/03 against annotations the servers themselves ship, including the honest-degradation path when nothing is declared. None of these services mints a credential with a local authorization rule to exercise, so every case ran with a NullProbe: authority and effectiveness stay unknown and EFF-06 SKIPs. Probe-backed authority/effectiveness classification (E2/E3) remains validated on the controlled testbed only. Runs are pinned to the exact package versions shown; they need the servers on the machine, so they are re-run by python experiments/case_studies.py, not by the CI reproduction gate.
@modelcontextprotocol/server-memory@2026.8.31
- All nine tools ship full annotations; the three readOnlyHint=true tools (read_graph, search_nodes, open_nodes) caused no observed change — EFF-01 verified against the server's own declarations.
- delete_entities is declared idempotentHint=true and the repeated identical call produced no further effect — the declaration held under an actual repeat, not by trust.
- Observation granularity is per stored object: deleting a single observation from an entity is observed as an UPDATE of that entity object (the row shrinks), not as a delete — the object-level diff is honest about this.
@modelcontextprotocol/server-filesystem@2026.8.31
- All ten readOnlyHint=true tools caused no observed filesystem change — EFF-01 verified against the server's own declarations.
- move_file is observed as create(destination) + delete(source); its destructiveHint=true declaration covers the observed delete (EFF-02).
- This call improved the instrument: EFF-02's first implementation read only the record's headline effect, where create outranks delete — so move_file's delete was invisible to it (a lying tool could mask a delete by also creating something). EFF-02 now reads per-target ops; the gap was found by this case study, not by the testbed.
- write_file and create_directory declare idempotentHint=true; the repeated identical calls left the tree byte-identical, so the claim held under an actual repeat.
@executeautomation/database-server@1.1.0
- No tool declares any annotation — the ecosystem's common case. EFF-01/03 SKIP (nothing declared to verify); the observed DELETE from write_query is covered by the spec's pessimistic default for an unset destructiveHint and is therefore consistent, not a contradiction.
- Effects are still attributed per row (create/update/delete of items/<id>) even with nothing declared — observation does not depend on annotations.
- append_insight turned out to persist into a table the audit never created (mcp_insights) — caught because the observer introspects sqlite_master rather than diffing a hand-listed table set. (Its archived Python predecessor kept insights in process memory; verifying before writing this note corrected our own assumption.)
- Observation granularity, honestly: create_table of an EMPTY table is a schema-level change below the row-level diff (no rows → no per-object delta); the table becomes visible the moment it holds a row. The record for create_table therefore shows no observed per-object change.
Method & reproduction
#Every number on this page is read from the experiment runners' JSON outputs, so the page cannot drift from what was measured. The observer reads the server's SQLite store directly; the probe replays the server's real authorization rule out-of-band (shared code, faithful by construction). Effects a channel cannot establish are reported unknown and never counted as evidence. Full methodology, limitations and related-work positioning: docs/effect-aware-conformance.md.
Reproduce from a clean state:
python experiments/run_all.py — regenerates the JSON, the markdown
tables and this page. Flagship effect-evidence reports:
python experiments/make_report.py.