mcp-proof / evaluation

Effect-aware conformance

Can declared MCP effects be verified against independently observed external state?

Declared MCP semantics can be misleading even when tool responses and tool names look correct. Each experiment compares a declared or inferred property against ground truth established outside the audited MCP surface — an out-of-band state observer and a probe that exercises created objects. E1 is the detection experiment; E2 validates that the persistence/authority distinction is operationally real (construct validation); E3 measures effectiveness after lifecycle events (residual authority); the case studies take the instrument to servers we did not build.

Testbed controlled · synthetic
Ground truth out-of-band state
Experiments 3 · 3 detectors each
Planted lies 6 (3 response-invisible)
Case studies 3 third-party servers

E1–E3 are controlled-testbed measurements of detection behaviour on planted inconsistencies — not production prevalence. Ground truth is read out-of-band from the state store, never through the audited MCP surface. The case studies run the same instrument against published third-party servers.

Artifacts: effect-report-honest.html effect-report-silent-keymint.html summary.md e1.json e2.json e3.json case-memory.json case-filesystem.json case-sqlite.json

E1

Declared effect vs observed effect

#
Research question

The detection experiment: can effect observation catch annotation lies that a response-only auditor and a name heuristic miss?

Finding

Only effect observation detects all three response-invisible lies. Each baseline reaches recall 0.333 with a structural blind spot the observation channel cannot fix.

DetectorPrecisionRecallF1
Effect-awareout-of-band state observation · TP 6 · FP 0 · FN 01.0001.0001.000
Response-leveltool responses only (current auditor) · TP 2 · FP 0 · FN 41.0000.3330.500
Name heuristictool names/descriptions only · TP 2 · FP 0 · FN 41.0000.3330.500

aggregated over the honest server + six single-lie variants; per-lie detections below

Planted lieLying toolEffect-awareResponse-levelName heuristic
effect visible in the response
lie-readonlysave_notecaughtcaughtmissed
hide-destructivedelete_notecaughtcaughtcaught
webhook-readonlycreate_webhookcaughtmissedcaught
effect invisible in the response — only observation can see it
silent-keymintget_notecaughtmissedmissed
shadow-webhooklist_notescaughtmissedmissed
phantom-writepingcaughtmissedmissed

● caught○ missed — highlighted rows are the three response-invisible lies

Inspect evidence — per-variant detections
Server variantGround truth (lying tool)Effect-aware flaggedResponse-level flaggedName heuristic flagged
honest————
lie-readonlysave_notesave_notesave_note—
hide-destructivedelete_notedelete_notedelete_notedelete_note
webhook-readonlycreate_webhookcreate_webhook—create_webhook
silent-keymintget_noteget_note——
shadow-webhooklist_noteslist_notes——
phantom-writepingping——

One audit per variant; every detector reads the same execution, so differences come from the observation channel, not from run-to-run noise. The honest variant measures false positives — no detector flagged anything on it.

Oracle: the testbed's mutation ledger — each variant plants exactly one annotation lie in one tool. Fairness: all three detectors read the same execution records; they differ only in which channel they may consult (external state, response text, or metadata).

E2

Persistence is not authority

#
Research question

Construct validation, not a benchmark: is the persistence/authority distinction operationally real — does probe-backed classification separate authority-bearing objects from ordinary persistent objects where “persistent ⇒ authority” and a credential-name keyword go wrong?

Finding

The probe classifies by use, not appearance — it is correct on the decoy note named api_key_backup (fools both baselines' signals) and on a credential that never persisted.

ClassifierAccuracyTPFPTNFN
Probe (exercise)attempts to use the object1.0005030
Name keywordcredential-looking names count as authority0.8755120
Persistence ⇒ authoritypersistent objects count as authority0.6254211

object corpus spans the persistent × authority quadrants; per-object classifications below

ObjectPersistentAuthority (truth)Persistence saysName saysProbe saysNote
api_keys/key_0001truetruetruetruetrue—
api_keys/key_0002truetruetruetruetrue—
webhooks/wh_0001truetruetruetruetrue—
share_links/share_0001truetruetruetruetrue—
notes/welcometruefalsetrue (disagrees with ground truth)falsefalse—
notes/api_key_backuptruefalsetrue (disagrees with ground truth)true (disagrees with ground truth)falsedecoy: credential-looking note name
notes/scratchfalsefalsefalsefalsefalseephemeral note (created, then deleted in-plan)
api_keys/key_0003falsetruefalse (disagrees with ground truth)truetruecounterexample: credential minted then removed out-of-band — authority without persistence

● true○ false prediction disagrees with ground truth

Oracle: by-construction authority labels — a table's objects are authority-bearing iff the testbed's real authorization rule accepts them as credentials. Probing time: each object is exercised at creation time, while it genuinely exists; persistence is judged against the final snapshot.

E3

Existence is not current effectiveness

#
Research question

A lifecycle measurement: after an event (grant revoked, key revoked, key deleted, TTL expired), does an exercise probe report an object's true effectiveness where “it is still listed” and “its grant is still active” do not — and why authorized_by cannot be read as depends_on?

Finding

In grant_revoked the key remains effective although the grant that authorized it is gone — residual authority. The delegation-centric view calls it dead (the one false-ineffective); the probe does not, and existence never distinguishes a live key from a listed-but-dead one.

DetectorAccuracyFalse-effectiveFalse-ineffective
Probe (exercise)attempts to use the object1.00000
Existence (inventory)listed ⇒ effective0.50030
Delegation-centricgrant active ⇒ effective0.33331 — residual authority missed

false-ineffective is the dangerous error: believing an object is neutralized while it still works

Lifecycle scenarioTruthExistence saysDelegation saysProbe saysNote
no_eventeffectiveeffectiveeffectiveeffectivekey minted under an active grant, nothing revoked
grant_revokedeffectiveeffectivedead (disagrees with ground truth)effectivethe authorizing grant is revoked; the key still works (residual authority) residual authority
key_revokeddeadeffective (disagrees with ground truth)effective (disagrees with ground truth)deadthe key itself is revoked; it is still listed but dead
key_unlisteddeaddeadeffective (disagrees with ground truth)deadthe key row is deleted; not listed and dead
ttl_expireddeadeffective (disagrees with ground truth)effective (disagrees with ground truth)deadthe key's TTL has passed; still listed but dead
grant_revoked_cascadedeadeffective (disagrees with ground truth)deaddeadgrant revoked on a server that cascades; the key is dead

● effective○ dead prediction disagrees with ground truth

Inspect evidence — per-scenario probe transcript
ScenarioKey rowGrantProbe statusProbe basis
no_eventlistedactiveeffectiveexercised api_keys.secret: server auth rule accepts it (status/expiry checked)
grant_revokedlistedrevokedeffectiveexercised api_keys.secret: server auth rule accepts it (status/expiry checked)
key_revokedlistedactiveineffectiveexercised api_keys.secret: server auth rule rejects it (status/expiry checked)
key_unlistedabsentactiveineffectiveobject not listed; nothing to exercise
ttl_expiredlistedactiveineffectiveexercised api_keys.secret: server auth rule rejects it (status/expiry checked)
grant_revoked_cascadelistedrevokedineffectiveexercised api_keys.secret: server auth rule rejects it (status/expiry/grant checked)

The probe basis is the evidence line the exercise probe recorded: which credential column it exercised and which rule accepted or rejected it. The key is api_keys/key_0001 in every scenario; each scenario runs on a fresh server.

Oracle: intended effectiveness per scenario, by construction — modelled on documented provider behaviour (a revoked grant leaving a minted key alive; keys listed after revocation). Event application: lifecycle events are applied out-of-band, never through the audited tools.

CS

Case studies — third-party servers

#
Research question

Does the instrument work on servers we did not build? 3 published MCP servers — official reference servers and a community server — across three store types (JSONL file, directory tree, SQLite) and two annotation profiles (fully annotated, none declared).

Finding

No annotation contradiction on any of the 3 audited servers — the honest result. The declarations that could be checked held under observation (readOnly tools wrote nothing; deletes were declared destructive; idempotent claims held under an actual repeat), and where nothing was declared the checks SKIPped instead of inventing a verdict.

Server (pinned)Store · observerAnnotationsObserved effectsEffect checks
@modelcontextprotocol/server-memory@2026.8.31modelcontextprotocol (official reference server)knowledge-graph JSONL file (MEMORY_FILE_PATH)an out-of-band parse of the JSONL store into per-entity/per-relation objects9/9 tools3 readOnly · 3 destructive · 6 idempotent2c / 2u / 2dover 10 calls
EFF-01 PASSEFF-02 PASSEFF-03 PASSEFF-06 SKIP
effect-evidence report · raw JSON
@modelcontextprotocol/server-filesystem@2026.8.31modelcontextprotocol (official reference server)jailed directory treethe stock FilesystemObserver over the jail root14/14 tools10 readOnly · 3 destructive · 2 idempotent3c / 1u / 0dover 16 calls
EFF-01 PASSEFF-02 PASSEFF-03 PASSEFF-06 SKIP
effect-evidence report · raw JSON
@executeautomation/database-server@1.1.0ExecuteAutomation (community)SQLite database filethe stock SqliteObservernone declared0 readOnly · 0 destructive · 0 idempotent3c / 1u / 1dover 10 calls
EFF-01 SKIPEFF-02 PASSEFF-03 SKIPEFF-06 SKIP
effect-evidence report · raw JSON

Scope of this validation: these runs exercise the observation half of the lane for real — out-of-band per-object effect attribution and EFF-01/02/03 against annotations the servers themselves ship, including the honest-degradation path when nothing is declared. None of these services mints a credential with a local authorization rule to exercise, so every case ran with a NullProbe: authority and effectiveness stay unknown and EFF-06 SKIPs. Probe-backed authority/effectiveness classification (E2/E3) remains validated on the controlled testbed only. Runs are pinned to the exact package versions shown; they need the servers on the machine, so they are re-run by python experiments/case_studies.py, not by the CI reproduction gate.

@modelcontextprotocol/server-memory@2026.8.31

  • All nine tools ship full annotations; the three readOnlyHint=true tools (read_graph, search_nodes, open_nodes) caused no observed change — EFF-01 verified against the server's own declarations.
  • delete_entities is declared idempotentHint=true and the repeated identical call produced no further effect — the declaration held under an actual repeat, not by trust.
  • Observation granularity is per stored object: deleting a single observation from an entity is observed as an UPDATE of that entity object (the row shrinks), not as a delete — the object-level diff is honest about this.

@modelcontextprotocol/server-filesystem@2026.8.31

  • All ten readOnlyHint=true tools caused no observed filesystem change — EFF-01 verified against the server's own declarations.
  • move_file is observed as create(destination) + delete(source); its destructiveHint=true declaration covers the observed delete (EFF-02).
  • This call improved the instrument: EFF-02's first implementation read only the record's headline effect, where create outranks delete — so move_file's delete was invisible to it (a lying tool could mask a delete by also creating something). EFF-02 now reads per-target ops; the gap was found by this case study, not by the testbed.
  • write_file and create_directory declare idempotentHint=true; the repeated identical calls left the tree byte-identical, so the claim held under an actual repeat.

@executeautomation/database-server@1.1.0

  • No tool declares any annotation — the ecosystem's common case. EFF-01/03 SKIP (nothing declared to verify); the observed DELETE from write_query is covered by the spec's pessimistic default for an unset destructiveHint and is therefore consistent, not a contradiction.
  • Effects are still attributed per row (create/update/delete of items/<id>) even with nothing declared — observation does not depend on annotations.
  • append_insight turned out to persist into a table the audit never created (mcp_insights) — caught because the observer introspects sqlite_master rather than diffing a hand-listed table set. (Its archived Python predecessor kept insights in process memory; verifying before writing this note corrected our own assumption.)
  • Observation granularity, honestly: create_table of an EMPTY table is a schema-level change below the row-level diff (no rows → no per-object delta); the table becomes visible the moment it holds a row. The record for create_table therefore shows no observed per-object change.

Method & reproduction

#

Every number on this page is read from the experiment runners' JSON outputs, so the page cannot drift from what was measured. The observer reads the server's SQLite store directly; the probe replays the server's real authorization rule out-of-band (shared code, faithful by construction). Effects a channel cannot establish are reported unknown and never counted as evidence. Full methodology, limitations and related-work positioning: docs/effect-aware-conformance.md.

Reproduce from a clean state: python experiments/run_all.py — regenerates the JSON, the markdown tables and this page. Flagship effect-evidence reports: python experiments/make_report.py.