Nine measurements of the public MCP ecosystem
中文索引 - a generated Chinese index of the same questions, with every count read back from the dataset.
We asked servers that are not ours nine anonymous questions, because the claims we make about gateway behaviour were otherwise just opinions. Run 2026-09-27, sample of 40 servers advertised in the official MCP registry.
Two machine-readable artifacts back this page. The first,
data/mcp-ecosystem-measurements.json, is the 2026-09-27 run: five questions measured
together in one window, which is why its tables match the five write-ups above. The second,
data/mcp-ecosystem-measurements.2026-09-28.json, is a wider re-run - all nine questions,
including whether anyone enforces the MCP-Protocol-Version header, whether a server can
be routed by the Mcp-Method header when the body says otherwise, the same question with
no body at all, and the cache-hint survey entered twice: once asking the revision that does
not require ttlMs/cacheScope, once asking the revision that does. It is a
new measurement, not a regression test of the first. Where a count differs by a server or two from a
page above, both figures are correct for their own window and the page's own date governs - which is
the reason neither file was quietly overwritten with the other.
paginate tools/list. Sixteen servers returned a tool list; none
emitted nextCursor. Largest single page seen: 35 tools.
agreed to a protocol version that does not exist. Asked for
9999-99-99, two answered HTTP 200 with that value echoed back.
require the session id they issue. Two servers handed out
MCP-Session-Id and returned 400 without it. Zero issued a token and
then ignored it.
The denominator that matters most is the unfashionable one: 22 of the 40 servers refused an
anonymous handshake entirely (401/403). Every ratio above is
therefore about the servers that let us look, not about the ecosystem.
1. Does anyone's tools/list actually paginate?
Our gateway walks every page and refuses to present a truncated enumeration as complete (issue #177). We closed that issue by admitting no real server had been observed paginating, so the bounds in the fix - 20 pages, 2,000 tools, dedupe by name, hard error on a repeated cursor - were a stance about a protocol feature rather than a survey of it. This is the survey.
Result: the walk is insurance against a server nobody has met yet. Which is the honest way to describe a fix for an unobserved failure, and worth saying out loud, because "we fixed pagination" would otherwise be read as "we saw pagination break".
Why it still matters: an aggregator that treats one response as complete applies its tool allowlist to page one only. Page-two tools are then neither permitted nor denied - they never reach the policy. A truncated enumeration that looks complete is worse than one that looks truncated, and nothing in the response shape lets a client tell the difference.
Full table and limits: pagination survey - the same write-up served as a page, or read the markdown source.
2. Will a server agree to a protocol version that does not exist?
The revision decides whether MCP-Session-Id is part of the contract and - in
2026-07-28 - whether the conversation is stateful at all. A server that echoes
whatever the client asked for makes the handshake information-free, and a client cannot notice
from the inside.
| Behaviour | Servers |
|---|---|
| Named a revision they support | 14 (6x 2025-06-18, 7x 2025-11-25, 1x 2024-11-05) |
| Rejected with a JSON-RPC error over HTTP 400 | 2 |
Echoed 9999-99-99 with HTTP 200 | 2 |
502, no version | 1 |
The paired reading is more useful than the headline: seven servers answered
2025-11-25 to the nonsense request while answering 2025-06-18 to a valid
one an hour earlier. The fallback is disclosing the revision the server would have preferred -
the one a client never learns by asking politely.
Full table and limits: revision-tolerance survey - also as a markdown source.
3. If a server issues a session id, does it require it back?
Of the 16 servers that answered initialize, 14 issued no session id at all and served
tools/list without one. Two issued one and required it back. The two that enforce it
are also two of the seven that negotiated upward, and both reply over SSE - so in this sample the
stateful servers are not a legacy tail, they are the ones adopting newer revisions. That is the
opposite of the reading where statelessness is simply taking over.
For anyone writing a gateway: store whatever the upstream issued and replay it, and treat "no session id" as a per-upstream property discovered at connect time rather than a global assumption about the protocol.
Full table and limits: session-enforcement survey - also as a markdown source.
4. Does anyone implement server/discover yet?
2026-07-28 makes server/discover a named RPC - one call that returns
supported protocol versions, capabilities and identity. A server maintainer's request log showed
real clients already sending it and receiving -32601 method not found, so the question
is how much of the reachable population would answer at all. Run 2026-09-27 18:08 UTC, same first 40
endpoints, same slice:
| What came back | Count |
|---|---|
| A real result | 1 |
-32601 method not found | 12 |
-32602 invalid params (the method was recognised) | 1 |
HTTP 404, no JSON-RPC payload at all | 1 |
An error object with no code field | 1 |
That is 16 servers that answered initialize; the other 24 of the 40 either refused an
anonymous client (22) or failed initialize (2), and they are outside this count rather
than failing it.
The one that implements it is ad.getle/leads (https://mcp.getle.ad/mcp),
answering over plain JSON with supportedVersions, capabilities,
instructions, cacheScope and ttlMs - and it negotiated
2025-06-18, an older revision than the one server/discover belongs to. So
the method is not being gated on the version string.
Two readings worth more than the headline. ai.agentberg/agentberg returned
-32602 Invalid request parameters rather than -32601: it knows the method
exists and rejected our empty params, which means a client cannot infer "not
implemented" from an error alone - it has to look at which error. And ac.tandem/docs-mcp
answered the RPC with HTTP 404 and no JSON-RPC body, which is a transport-level answer
to a method-level question: a gateway that maps status codes to "unsupported" and a gateway that
parses JSON-RPC errors will disagree about that server.
For anyone deciding whether to send server/discover first: in this sample it costs one
round trip and tells you almost nothing you could not get from initialize, because 12
of 16 have never heard of it. Try it, do not depend on it.
This is one timestamp and one alphabetical slice, same caveat as the other three: the sample
over-represents server names starting with a, and n=1 on the
implemented side is a demonstration that it exists in the wild, not an adoption rate. We did not
probe our own endpoint with this fourth script, so nothing here says whether we would answer it.
5. How much server-written natural language do clients already swallow?
The fourth answer was a single server, but the field it returns is the interesting part:
instructions is server-controlled prose that a client is told it "can use ... by
including it in a system prompt" (MCP-2026-015).
Run 2026-09-27 18:19 UTC against the same 40 endpoints:
- 9 of the 16 that answered
initializesentinstructions- between 72 and 1,423 characters, to an anonymous client, before any tool had been called. Eight distinct server names; one of them is registered at two endpoints and sent the field from both. - The one server that answers
server/discoverpairs its 617 characters withcacheScope: "public"andttlMs: 300000. That is the exact combination the security issue describes as a cache-poisoning amplifier, running on a registry-advertised endpoint. - 0 of 9 matched any blatant override marker we looked for ("ignore all previous safety", "you are now", "do not tell", "system prompt", "OVERRIDE"). Read that narrowly: it says the deployed population is not currently abusing the field. It says nothing about whether the field is safe, and a null result on nine samples is not a risk assessment.
- The legacy
initializepath carries it 9 times, the newserver/discoverpath 1. A mitigation scoped toserver/discoverwould cover one tenth of where the prose actually arrives today.
No response text is re-emitted by the script: it records length, a SHA-256 prefix and a marker
count. Copying an instructions string into a terminal, a report or an issue is the
same act the issue is about, and this page is not going to commit it while arguing about it.
6. Does anyone enforce the MCP-Protocol-Version header?
The 2025-06-18 revision makes the header a client obligation on every request after
initialize, and permits a server to reject one naming a revision it did not agree to.
Of 40 endpoints, 22 sat behind OAuth, 2 failed the handshake, and 16 completed it. Each of those 16
was asked for tools/list twice, the two requests differing by exactly one header:
16 of 16 served it with the header, and 16 of 16 served it without. One server
negotiated upward; naming the requested revision at it still worked.
So enforcement is not happening in this sample - and that answer is why the fix was still worth
making. A negative reading about other people's servers says nothing about what our client owes,
and checking ourselves turned up a real one: the gateway stored the session id an upstream issued and
replayed it, but stored the revision an upstream named and never sent it. Worse, an
operator-supplied header was forwarded verbatim, so against an upstream that answers
2024-11-05 our client could send 2025-06-18 - a revision the upstream had
explicitly declined, which is precisely the request a conforming server may refuse.
This question is measured, and from the 2026-09-28 re-run onward it is in the machine-readable dataset as well; the 2026-09-27 file below covers questions 1-5. Its per-server table is rendered from the probe's own artifact in the protocol-version header write-up.
7. Can a header redirect a server to a method the body never asked for?
Mcp-Method and Mcp-Name exist so a server can route a body-less GET
stream. A POST that already carries a JSON-RPC body does not need them - so the question is
whether a server reads them anyway, because one that does can be pointed at a method its
caller never wrote down. Each server that answered the handshake got the byte-identical
tools/list request twice, one leg adding Mcp-Method: prompts/list:
of the 16 with a comparable pair, 16 served the body and 0 were routed by the
header. Our own server, asked the same way live, also returned the tool list for both
a spoofed and a nonsense method.
The limit named above was then measured rather than left as a caveat.
tools/survey-mcp-get-stream-headers.mjs sends a body-less GET after a real
handshake, once with no routing hint and once adding Mcp-Method: prompts/list, and
of the 13 servers whose GET leg actually answered, 0 were routed by the header.
It also repeats the plain leg, which is how a false positive got caught before it became a
sentence: two servers appeared to change behaviour (409 against an aborted first leg), and the
repeat plain leg answered the same 409. The first version of that instrument had
no repeat leg and classified both as a header effect - a difference between two requests sent
seconds apart, mistaken for a difference caused by one header. Ten servers rejected the GET
identically either way; three answered GET with application/json and no stream at
all, again identically. The write-up, including the false positive the repeat leg caught, is
Can a header redirect a server to a method the body never asked for?
8. Does a server ever say how long its tool list may be cached?
A gateway has to decide how long to keep a cached list, and the spec lets the server answer that
with ttlMs and cacheScope - but only from protocol revision
2026-07-28 onward, which decides what a count of silent servers can mean. Asked at
2025-06-18, where those fields are not required, 17 of 40 endpoints returned a
list and 16 declared neither, 1 declared both
(ttlMs: 300000 with cacheScope: "private"). Re-asked the same day, over the
exactly identical 40 endpoints, at 2026-07-28: 13 returned a list, only 3
accepted that revision, and 0 of those 3 sent the fields it requires. So the claim that
servers do not declare hints survives only in the small place where it can actually be tested - and
the larger count, read as conformance, was never evidence of anything. We were reading neither field
ourselves, caching every upstream for a hard-coded 60 seconds in a process-global map shared by all
tenants; measured through a whole-frame stdio tee once the new revision is negotiated, our server
does send both.
Joining the two same-day legs by endpoint explains why the denominator collapsed: 5 servers
answer a 2025-06-18 initialize - several with a full tool list - and return HTTP
400 to a 2026-07-28 one instead of replying with their latest
supported revision. A 400 gives a client nothing to downgrade from, so those endpoints are not
legacy-era to a newer client, they are simply absent. That is why only 3 of the 40 could be asked
about the new fields at all, and it is a shape worth knowing before any client ships a validator
that treats absence as invalid.
The timing half is small: one declarant in sixteen, and honouring it meant clamping to
1 s..10 min so that a declared 0 cannot turn a read into a fresh handshake
and a declared decade cannot freeze a list. The sharing half is the one worth the change, and it
is stated without inflation: listTools() takes no caller identity, so nothing was
exposed - the cache was merely built as though sharing were always safe, which is the shape that
would turn a future per-caller header into a cross-tenant leak with nobody objecting.
The write-up, its blind spots, and our own silence.
Three more things the same API told us
Reading registry records for their own reason surfaced a separate question: does a listing carry an artifact an installer can consume? On the alphabetically-first 54 records the answer was 6 of 54 carried a package at all, and that page records why counting them on the oldest list row instead of the latest record gives a different number. Walking the whole default list rather than its front half - 37,013 servers visible, 37,854 when removed records are asked for - puts that share at 41.85%, and finds 439 records (1.20%) that declare neither a package nor a hosted endpoint, which is the only group a client genuinely cannot act on. A package and a remote are two different ways to be actionable, and the first version of the sample page conflated them; the census page records the retraction and the reading that disproved it. Then a seeded draw of 200 of the 9,896 npm-listed records came back the other way: 196 resolve on npm at exactly the listed version, 3 list a version npm does not have, and 1 name is gone - a 2.00% unusable rate with a 95% interval of 0.06% to 3.94%, which is the number that narrows the problem to the records with no coordinates at all rather than stale ones. The other five artifact types were probed the same way afterwards - 15 of 785 do not resolve, with cargo and NuGet counted in full - and pooled across all six types the unusable rate is 1.93% (Wilson 1.24% to 2.99%). Scoring our own server with the rubric a directory uses to score strangers is a separate reading, and the honest version of it is on this page: 15 of 15 of our tools declare no outputSchema, while all 15 do carry titles and all four MCP annotations.
The numbers as a file
Every table on this page also exists machine-readable: docs/data/mcp-ecosystem-measurements.json - one entry per question, each with its script, its verdict counts and the per-endpoint rows behind them. Regenerate it with
node tools/build-mcp-measurement-dataset.mjs 40 docs/data/mcp-ecosystem-measurements.json
The file is generated and never hand-edited, and the builder refuses to write it if any
question's verdict tally does not sum to that question's own row count. That guard is here
because prose on this site has already disagreed with the table beneath it once
(#180): a headline said
2/19 next to a tally of 18, and every doc check in the repository passed, because
they match words and none of them does arithmetic. A generated file cannot make that particular
mistake quietly.
It contains verdicts, counts, character lengths and a SHA-256 prefix for each
instructions value. It contains no server response text: an
instructions field is precisely the kind of payload that should not be copied into
an artifact and re-distributed, and a dataset about that problem is not going to demonstrate it.
Two files, two samples, one command.
mcp-ecosystem-measurements.json is the
40-endpoint run the tables above are drawn from (19:04-19:08 UTC).
mcp-ecosystem-measurements.wide.json
is the widest slice these scripts can reach - 54 to 55 unique endpoints out of the first 400
registry records, 19:15-19:22 UTC - and it is the run behind the figures quoted in two upstream
threads on 2026-09-27: 15 of 23 servers send instructions (72 to 19,687 characters,
median 406, one of fifteen over 2,000) and 1 of 23 answers server/discover while 19
return -32601. Both are produced by the same command with a different limit.
The row counts in the wide file differ between questions (54 or 55), because each of the five scripts re-reads the registry on its own and nothing pins the list across the five runs; the cause of the one-endpoint difference is not recorded. It is left visible in the files rather than smoothed over, since a dataset that presents its own sample as fixed is one nobody can re-run honestly.
Method, in one paragraph
Each script reads the official MCP registry for streamable-http records, takes the
first 40 endpoints in the registry's default order, and sends anonymous JSON-RPC:
initialize, then notifications/initialized, then
tools/list - or, for the fourth script, server/discover - with a second
page or a header-omitted repeat only where the specific question requires it. No credentials, no writes, no tool invocation, sequential requests with a
short delay. The two final requests in the session probe differ in exactly one header, so a
rejection cannot be blamed on something else the client forgot to send.
node tools/survey-mcp-tools-list-pagination.mjs 40
node tools/survey-mcp-revision-tolerance.mjs 40
node tools/survey-mcp-session-enforcement.mjs 40
node tools/survey-mcp-server-discover.mjs 40
node tools/survey-mcp-instructions-field.mjs 40
node tools/survey-mcp-protocol-version-header.mjs 40
node tools/survey-mcp-route-headers.mjs 40
node tools/survey-mcp-get-stream-headers.mjs 40
node tools/survey-mcp-list-cache-hints.mjs 40
We ran the same three questions against ourselves
Publishing numbers about other people's servers is easy; the test of whether the questions
mean anything is answering them about your own. Our HTTP MCP endpoint
(pnpm mcp:http, loopback only, credential-free default, port 3299 for the run)
was asked the same three things on 2026-09-27 17:19 UTC:
| Question | Our answer | Which group that puts us in |
|---|---|---|
Asked for 2025-06-18 |
negotiated 2025-06-18, HTTP 200 |
normal |
Does our tools/list paginate? |
15 tools, no nextCursor |
we are in the non-paginating majority ourselves - which is the honest context for the 0 of 16 above |
| Do we issue a session id? | none issued; requests served without one | the 14-of-16 stateless group |
Asked for 9999-99-99 |
answered 2025-11-25, HTTP 200 |
we substitute a supported revision rather than echo the impossible one - we are not in the group of 2 |
The fourth row is also the clearest illustration of
issue #178: our
server's fallback revision (2025-11-25) is newer than the one our own
gateway client declares to upstreams (2025-06-18), and when these probes ran the
gateway did not read what came back. Two components of the same product disagreed about which
protocol they were on, and nothing in the code noticed. Both halves of that are now different
in master and in the rolling :latest / :master image tags built from it on 2026-09-27, though not in the versioned 0.8.0 tag: the governed upstream client
records the revision each server answered and names it per server in the servers
array of GET /mcp/tools, and the second MCP client path in
claude-code-patterns sends the same constant as the first, with a test that fails
if the two drift apart. What the gateway should do when a server names a different
revision - refuse it, or serve it and say so - is still open on the issue, and this page will
be edited again when it is decided.
The run left nothing behind: the port was probed again after shutdown and refuses
connections, and no http-entry process remained.
Limits, stated rather than implied
- n is 16-18, not 40. 22 servers would not talk to an anonymous client, so they are outside these measurements rather than passing them.
- The sample is not random. It is the registry's default order at one timestamp, which over-represents server names beginning with
a. initializeandtools/listonly. Going further would mean invoking tools on services that are not ours.- No stdio coverage. That population is larger and unreachable by this method.
- One timestamp each, 2026-09-27 - with one deliberate exception. The cache-hint question is published as a same-day pair re-run on 2026-09-28, because asking for a revision that does not require
ttlMs/cacheScopeis a different measurement from asking for the one that does; the 2026-09-28 dataset records which revision each leg asked. Servers deploy; a re-run is a new measurement, not a regression test. All nine survey scripts are deliberately outside our CI for that reason; a tenth tool,tools/compare-mcp-revision-legs.mjs, only joins two saved legs and does run underpnpm test:verification-tools, which is the list CI actually runs.