Explore · experiment
I Scanned 16 MCP Servers for Safety Hints. The Scanner Flagged Every One.
Published 20 May, 2026 · experiment
On 2026-05-21 I scanned 16 public MCP servers for tool annotations, the hints agent clients use to decide when a human approves an action. The scanner flagged all 16, and one of those results is wrong. Here is what the gap means for approval gates, and what to check before you adopt a server.
- Question
- Do public community MCP servers declare the tool annotations that agent clients use to decide when a human must approve an action?
- Activity
- complete
I Scanned 16 MCP Servers for Safety Hints. The Scanner Flagged Every One.
On May 21, 2026, I ran an MCP scanner I’m building against a batch of public community servers. The scanner runs several checks. This piece is about one of them: does each tool say whether it only reads, whether it can destroy data, and whether it reaches outside its own environment? I expected some servers to pass that check and some to fail. Instead, that one check flagged every server in the batch.
If that result is even roughly right, the gap is across the ecosystem, not in a few careless servers: most tools don’t carry the metadata that agent approval gates are built to read. One result in the batch doesn’t hold up against the server’s public documentation. That limits how much the rest of the result can prove, and I cover it below instead of hiding it.
Why this matters: MCP, the Model Context Protocol, is the open standard AI agents use to connect to outside tools such as file systems, browsers, databases and SaaS APIs. Agent clients (the app running the agent) can use tool annotations to decide which calls run on their own and which stop for a human. When those hints are missing, the client can’t tell list_directory from delete_record. The team that approves agent actions is left approving everything by hand or allowlisting everything, and anyone who wires community MCP servers into an agent workflow inherits that choice.
What a tool annotation tells the client
A tool annotation is a small block of metadata on each tool that describes how the tool behaves. It doesn’t describe what the tool does. Annotations weren’t in the first MCP spec revision (2024-11-05). They were added in revision 2025-03-26, and the MCP project’s own write-up on annotations lists four hints and their defaults:
| Hint | What it signals | Default if absent |
|---|---|---|
readOnlyHint | The tool doesn’t change its environment | false (assume it writes) |
destructiveHint | A write may be destructive rather than additive | true (assume it destroys) |
idempotentHint | Repeating the call with the same arguments has no further effect | false |
openWorldHint | The tool interacts with an open set of outside entities, such as the web | true |
Two details in the spec shape everything that follows. First, the defaults are pessimistic: a tool with no annotations should be treated as one that may write, may destroy and may reach the open world. Second, the spec’s tools page says clients must treat annotations as untrusted unless they come from trusted servers.
Annotations aren’t a security control. They tell a client how to route calls from servers you already trust. A malicious server can label its delete tool read-only. What annotations do is let an honest server tell your approval gate which of its tools deserve a human.
The scanner flagged all 16 servers
The scanner connects to each server, lists its tools and checks each tool definition for any of the hints above. The original batch had 17 servers. I’ve removed one because it’s my own homelab server and doesn’t belong in a community sample, which leaves 16:
| Server | What it does | Scanner result (2026-05-21) |
|---|---|---|
| browsermcp | Browser automation | No annotations found |
| calculator | Arithmetic | No annotations found |
| chrome-devtools | Browser debugging | No annotations found |
| context7-mcp | Documentation retrieval | No annotations found |
| datetime-server | Date and time | No annotations found |
| deepwiki-mcp | Repository knowledge | No annotations found |
| everything-server | MCP reference/test server | No annotations found |
| filesystem | File operations | Flagged, but contradicted by its README (see below) |
| markdown-mcp | Document tools | No annotations found |
| memory-server | Persistent memory | No annotations found |
| onetrust-mcp | SaaS API bridge | No annotations found |
| playwright-mcp | Browser automation | No annotations found |
| puppeteer-mcp | Browser automation / scraping | No annotations found |
| sequential-thinking | Reasoning scaffold | No annotations found |
| sqlite-mcp | Database operations | No annotations found |
| youtube-transcript | Video transcript API | No annotations found |
Some cases were easy to read. browsermcp exposes browser_navigate, browser_go_back, browser_go_forward and browser_snapshot, and the scanner found no hints on any of them. Calculator’s add, subtract, multiply and divide also have none. That carries little risk on its own, but it shows how normal leaving the block out has become. For sqlite-mcp, a read query and a delete look the same to any client that connects.
The absence matters most on the servers that can change things: the browser automation servers, the database server and the API bridge.
One result I can’t stand behind
The official filesystem server’s README documents annotations for its tools. For example, it marks read_text_file and list_directory as read-only and uses destructiveHint and idempotentHint on the write operations. The scanner said it found nothing.
I have two explanations, and I haven’t confirmed either. The scanner may have tested an older build. Or it may look for the hints in the wrong place. In the TypeScript SDK, hints live in a separate annotations object on the tool. They aren’t top-level tool properties. My first draft of this post showed them at the top level, and a scanner written with the same wrong assumption would miss every correctly annotated tool.
That second explanation is a hypothesis, but it’s serious enough to change how the table should be read. Each entry means “the scanner found none”, not “confirmed absent”. Each server ran through ten iterations, and every iteration gave the same answer. For a check this deterministic, that proves the scanner is consistent. It doesn’t prove the scanner is right.
Why missing hints break approval gates
This is where I’d change my own first framing. I originally assumed that clients quietly allow unannotated tools. The spec’s defaults point the other way: a client that follows them treats every unannotated tool as a possibly destructive, open-world write. I have no survey of how clients actually behave, and behavior varies. Some clients ask for approval on every tool call whether it’s annotated or not.
The failure isn’t that unannotated tools run unchecked. It’s that every tool gets identical treatment. That leaves an operator three options:
- Prompt on every call. Nothing slips past, but the approver sees a stream of prompts that look the same.
- Allowlist the server. The agent runs, but reads and deletes now share one approval decision.
- Block the server. Safe, but the team loses the capability and often works around the block.
In the agent setups I’ve run and watched others run, option 1 wears down fastest. When every prompt looks the same, people approve without reading, and within weeks the gate is a formality. Then option 2 becomes the quiet default. Annotations are what let a gate ask about the delete and skip the directory listing.
Does any framework require annotations?
Not that I’ve found. No regulation or security framework I’ve read names MCP tool annotations. What exists are broader obligations that annotations can help you meet:
- EU AI Act, Article 14: high-risk AI systems must be designed and developed so that people can effectively oversee them while they’re in use.
- EU AI Act, Article 12: high-risk AI systems must technically allow events to be recorded automatically (logs) over their lifetime.
- OWASP Top 10 for Agentic Applications: ASI02, Tool Misuse and Exploitation, covers agents using legitimate tools in unsafe or unintended ways.
The link to annotations is my reading, not the text of either document. Oversight that treats a read and a delete identically is hard to call effective, and annotations are one input that lets you target it. Whether your agent workflows fall within Article 14’s scope is a question for your own legal review, not this post.
What’s still unresolved
- Scanner accuracy. Until I’ve rechecked the filesystem result, every “none found” in the table carries that doubt.
- Sample. Sixteen servers, picked by hand, on one date. It suggests a pattern. It doesn’t measure the ecosystem.
- Client behavior. I haven’t tested how two or three real agent clients handle unannotated tools against the spec’s defaults. That test decides how much this gap costs in practice.
- Trust. Even with full adoption, annotations from servers you don’t trust are worth nothing. The gap they close is narrower than “agent safety”.
- Movement. This is a snapshot as of June 6, 2026. The follow-up is to re-run the same servers, with the scanner fixed, and see whether they now carry annotations.
What I’d do next
Good (this week): For every MCP server your agents connect to, open the tool list and look for an annotations block on each tool. Don’t rely on a scanner’s summary, mine included. Write down which tools can write, delete or reach the internet, whether they say so or not.
Better (this quarter): Decide in writing how your agent clients treat unannotated tools, and base the decision on your business, not on whatever the client happens to do. My view is that unannotated tools from servers you haven’t reviewed should be treated as the spec’s worst-case defaults. Where do reads end and changes that need a human begin? That depends on what the tool touches in your environment. If your team writes MCP servers, add the hints; the SDK makes it a few lines:
server.registerTool(
"delete_record",
{
description: "Delete a database record by id",
inputSchema: { id: z.string() },
annotations: { readOnlyHint: false, destructiveHint: true, openWorldHint: false },
},
async ({ id }) => deleteRecord(id)
);
Best (this half): Make annotation review part of how you adopt a server, alongside the checks you’d already run on any third-party dependency. Only trust a server’s hints after you’ve reviewed the server. Then track how often approvers see prompts and how often they reject them. That tells you whether your gate still gets read or has become a formality. The right level depends on your environment, not on anyone’s benchmark.
If you’ve tested how a specific agent client handles unannotated tools, tell me what you found. That’s the piece of this I most want to get right.
Experiment run and last scanned 2026-05-21. One batch run of an MCP scanner I’m building, against 16 public community MCP servers (17 originally; one personal server removed), with ten iterations per server. The scanner runs several checks; the one reported here looked for readOnlyHint, destructiveHint, idempotentHint or openWorldHint on each tool definition. One result (filesystem) is contradicted by the server’s public documentation and is unresolved.
David O’Neil is a CISO and builder who runs AI agents on real projects and scans the tooling underneath them before trusting it.
ai agents · agent governance · supply chain
What would move this forward
experiment evidence (0 of 4 met)
-
A second dated re-run, e.g. the quarterly one the draft proposes, showing whether adoption moved (open)
-
A larger and clearly defined sample of public community servers, excluding personal or internal ones (open)
-
A check of how two or three real agent clients actually handle unannotated tools, given the spec's default hint values (open)
-
A corrected, runnable SDK annotation example (open)