EVIDENCE · SELF-ASSESSED, BOUNDED, RE-RUNNABLE
MCP security boundaries: what we have actually checked.
This page exists because the security evidence in this repository lived only as a GitHub file directory, where a search result cannot land on it. Every number below is produced by a command in this repository, and the first thing each section states is where that method stops. Nothing here is a production certification, and none of it is a claim of universal safety.
中文版本:MCP 安全边界 · Credential-free evidence: what one run proves about provider calls · Security evidence directory · What a published MCP image exposes
Every figure on this page is derived rather than typed: 23 probes counted from
tools/security-attack-regression.mjs, and 15 tool names counted from
packages/mcp-server/src/server.js. tools/public-repo-check.mjs asserts
both counts against those two files, so adding an attack case or a tool without updating this page
turns the build red instead of leaving the copy quietly wrong. If the sentence above ever stops
matching the code, the code is the authority.
The attack regression: twenty-three probes, all defended
tools/security-attack-regression.mjs starts a real gateway process and drives
twenty-three attack cases at it: cross-tenant isolation, header and identity forgery, role
checks, quota and rate behaviour, revocation, information leaks, and the MCP boundary. Each case
prints DEFENDED or BREACH, and any breach exits non-zero.
Run it yourself:
node tools/security-attack-regression.mjs
The committed page Security drill evidence is generated by executing that script and parsing its output. It is written only if exactly twenty-three verdicts are parsed and none is a breach, and it prints the count from the parsed verdicts rather than from a number typed into a template. If you add a twenty-fourth case and forget to update anything, the generator refuses to write instead of publishing a stale count.
Where this method stops. It is a self-assessed regression written by the people who wrote the code. It is not a penetration test by an independent party, and it is not an audit. It runs against the source build, so a green run does not certify the container image published to the registry - that is a separate question, handled below.
Two published-image reviews, by two different methods
These are not the same check and they do not cover the same ground. Reading them as one number is the mistake this section exists to prevent.
| Image | Method | What it required | What it concluded |
|---|---|---|---|
0.8.0 |
Read the layer tarballs straight out of the registry | No Docker engine; nothing is executed | Non-root node user, 8 setuid and 3 setgid files - and four of the eight
native modules in the linux/arm64 tag are x86-64, so the paths that load them
fail on that architecture |
0.4.9 |
Export a container filesystem with Docker | A working Docker engine | Default root user, 11 base-image SUID/SGID files; arm64 modules genuinely AArch64. This is the review the agent skill's pinned digest rests on |
Re-read either one without Docker:
node tools/inspect-image-filesystem.mjs 0.8.0
node tools/verify-image-roster.mjs 0.8.0
The second command answers a different question - which tool names the published image actually
exposes. It reports 15 for 0.8.0 and 9 for
0.4.9, read out of the image bytes rather than from any document, which is why the
pinned digest in the install instructions matters: nine of the fifteen names do not exist in that
older image.
Why the newest review is not reassurance. The 0.8.0 review is the one
that found the architecture regression. A newer document here means more was checked, not
that less is wrong.
Versions 0.5.0, 0.6.0 and 0.7.0 have no content review on
file. 0.4.8, 0.4.0 and 0.3.2 do, and are linked from the
security index.
What this page does not prove
- Not independent verification. Both the attack regression and the image reviews are ours. No third party has audited this project.
-
Not a runtime audit of the published container. The tarball review reads
bytes; it never starts the image. The
arm64finding is an ELF-header reading, so the failure mode is inferred from the binaries inside the tag, not observed in a running container on that architecture. - Not a population claim. Twenty-three cases are what we thought to write. They do not bound what an attacker will try.
- No production readiness, no recovery objectives. Nothing here is RTO or RPO evidence, and nothing here says the gateway is safe under your workload, your provider, or your network.
- Digest trust is delegated. Reading an image by digest assumes the registry serves the object it advertised. The review recomputes every blob digest it downloads, which catches corruption and substitution in transit, not a malicious registry that was never honest about the tag in the first place.
If you find something
Report privately through the security policy rather than an issue. If you re-run a command above and it disagrees with this page, that is itself worth reporting: the numbers here are derived, so a mismatch means either the artifact or the derivation moved, and we would rather find that out than defend a stale figure.