Nine measurements of the public MCP ecosystem

中文索引 - a generated Chinese index of the same questions, with every count read back from the dataset.

We asked servers that are not ours nine anonymous questions, because the claims we make about gateway behaviour were otherwise just opinions. Run 2026-09-27, sample of 40 servers advertised in the official MCP registry.

Two machine-readable artifacts back this page. The first, data/mcp-ecosystem-measurements.json, is the 2026-09-27 run: five questions measured together in one window, which is why its tables match the five write-ups above. The second, data/mcp-ecosystem-measurements.2026-09-28.json, is a wider re-run - all nine questions, including whether anyone enforces the MCP-Protocol-Version header, whether a server can be routed by the Mcp-Method header when the body says otherwise, the same question with no body at all, and the cache-hint survey entered twice: once asking the revision that does not require ttlMs/cacheScope, once asking the revision that does. It is a new measurement, not a regression test of the first. Where a count differs by a server or two from a page above, both figures are correct for their own window and the page's own date governs - which is the reason neither file was quietly overwritten with the other.

0 / 16

paginate tools/list. Sixteen servers returned a tool list; none emitted nextCursor. Largest single page seen: 35 tools.

2 / 19

agreed to a protocol version that does not exist. Asked for 9999-99-99, two answered HTTP 200 with that value echoed back.

2 of 2

require the session id they issue. Two servers handed out MCP-Session-Id and returned 400 without it. Zero issued a token and then ignored it.

The denominator that matters most is the unfashionable one: 22 of the 40 servers refused an anonymous handshake entirely (401/403). Every ratio above is therefore about the servers that let us look, not about the ecosystem.

1. Does anyone's tools/list actually paginate?

Our gateway walks every page and refuses to present a truncated enumeration as complete (issue #177). We closed that issue by admitting no real server had been observed paginating, so the bounds in the fix - 20 pages, 2,000 tools, dedupe by name, hard error on a repeated cursor - were a stance about a protocol feature rather than a survey of it. This is the survey.

Result: the walk is insurance against a server nobody has met yet. Which is the honest way to describe a fix for an unobserved failure, and worth saying out loud, because "we fixed pagination" would otherwise be read as "we saw pagination break".

Why it still matters: an aggregator that treats one response as complete applies its tool allowlist to page one only. Page-two tools are then neither permitted nor denied - they never reach the policy. A truncated enumeration that looks complete is worse than one that looks truncated, and nothing in the response shape lets a client tell the difference.

Full table and limits: pagination survey - the same write-up served as a page, or read the markdown source.

2. Will a server agree to a protocol version that does not exist?

The revision decides whether MCP-Session-Id is part of the contract and - in 2026-07-28 - whether the conversation is stateful at all. A server that echoes whatever the client asked for makes the handshake information-free, and a client cannot notice from the inside.

Answered initialize with protocolVersion: "9999-99-99"
BehaviourServers
Named a revision they support14 (6x 2025-06-18, 7x 2025-11-25, 1x 2024-11-05)
Rejected with a JSON-RPC error over HTTP 4002
Echoed 9999-99-99 with HTTP 2002
502, no version1

The paired reading is more useful than the headline: seven servers answered 2025-11-25 to the nonsense request while answering 2025-06-18 to a valid one an hour earlier. The fallback is disclosing the revision the server would have preferred - the one a client never learns by asking politely.

Full table and limits: revision-tolerance survey - also as a markdown source.

3. If a server issues a session id, does it require it back?

Of the 16 servers that answered initialize, 14 issued no session id at all and served tools/list without one. Two issued one and required it back. The two that enforce it are also two of the seven that negotiated upward, and both reply over SSE - so in this sample the stateful servers are not a legacy tail, they are the ones adopting newer revisions. That is the opposite of the reading where statelessness is simply taking over.

For anyone writing a gateway: store whatever the upstream issued and replay it, and treat "no session id" as a per-upstream property discovered at connect time rather than a global assumption about the protocol.

Full table and limits: session-enforcement survey - also as a markdown source.

4. Does anyone implement server/discover yet?

2026-07-28 makes server/discover a named RPC - one call that returns supported protocol versions, capabilities and identity. A server maintainer's request log showed real clients already sending it and receiving -32601 method not found, so the question is how much of the reachable population would answer at all. Run 2026-09-27 18:08 UTC, same first 40 endpoints, same slice:

What came backCount
A real result1
-32601 method not found12
-32602 invalid params (the method was recognised)1
HTTP 404, no JSON-RPC payload at all1
An error object with no code field1

That is 16 servers that answered initialize; the other 24 of the 40 either refused an anonymous client (22) or failed initialize (2), and they are outside this count rather than failing it.

The one that implements it is ad.getle/leads (https://mcp.getle.ad/mcp), answering over plain JSON with supportedVersions, capabilities, instructions, cacheScope and ttlMs - and it negotiated 2025-06-18, an older revision than the one server/discover belongs to. So the method is not being gated on the version string.

Two readings worth more than the headline. ai.agentberg/agentberg returned -32602 Invalid request parameters rather than -32601: it knows the method exists and rejected our empty params, which means a client cannot infer "not implemented" from an error alone - it has to look at which error. And ac.tandem/docs-mcp answered the RPC with HTTP 404 and no JSON-RPC body, which is a transport-level answer to a method-level question: a gateway that maps status codes to "unsupported" and a gateway that parses JSON-RPC errors will disagree about that server.

For anyone deciding whether to send server/discover first: in this sample it costs one round trip and tells you almost nothing you could not get from initialize, because 12 of 16 have never heard of it. Try it, do not depend on it.

This is one timestamp and one alphabetical slice, same caveat as the other three: the sample over-represents server names starting with a, and n=1 on the implemented side is a demonstration that it exists in the wild, not an adoption rate. We did not probe our own endpoint with this fourth script, so nothing here says whether we would answer it.

5. How much server-written natural language do clients already swallow?

The fourth answer was a single server, but the field it returns is the interesting part: instructions is server-controlled prose that a client is told it "can use ... by including it in a system prompt" (MCP-2026-015). Run 2026-09-27 18:19 UTC against the same 40 endpoints:

No response text is re-emitted by the script: it records length, a SHA-256 prefix and a marker count. Copying an instructions string into a terminal, a report or an issue is the same act the issue is about, and this page is not going to commit it while arguing about it.

6. Does anyone enforce the MCP-Protocol-Version header?

The 2025-06-18 revision makes the header a client obligation on every request after initialize, and permits a server to reject one naming a revision it did not agree to. Of 40 endpoints, 22 sat behind OAuth, 2 failed the handshake, and 16 completed it. Each of those 16 was asked for tools/list twice, the two requests differing by exactly one header: 16 of 16 served it with the header, and 16 of 16 served it without. One server negotiated upward; naming the requested revision at it still worked.

So enforcement is not happening in this sample - and that answer is why the fix was still worth making. A negative reading about other people's servers says nothing about what our client owes, and checking ourselves turned up a real one: the gateway stored the session id an upstream issued and replayed it, but stored the revision an upstream named and never sent it. Worse, an operator-supplied header was forwarded verbatim, so against an upstream that answers 2024-11-05 our client could send 2025-06-18 - a revision the upstream had explicitly declined, which is precisely the request a conforming server may refuse.

This question is measured, and from the 2026-09-28 re-run onward it is in the machine-readable dataset as well; the 2026-09-27 file below covers questions 1-5. Its per-server table is rendered from the probe's own artifact in the protocol-version header write-up.

7. Can a header redirect a server to a method the body never asked for?

Mcp-Method and Mcp-Name exist so a server can route a body-less GET stream. A POST that already carries a JSON-RPC body does not need them - so the question is whether a server reads them anyway, because one that does can be pointed at a method its caller never wrote down. Each server that answered the handshake got the byte-identical tools/list request twice, one leg adding Mcp-Method: prompts/list: of the 16 with a comparable pair, 16 served the body and 0 were routed by the header. Our own server, asked the same way live, also returned the tool list for both a spoofed and a nonsense method.

The limit named above was then measured rather than left as a caveat. tools/survey-mcp-get-stream-headers.mjs sends a body-less GET after a real handshake, once with no routing hint and once adding Mcp-Method: prompts/list, and of the 13 servers whose GET leg actually answered, 0 were routed by the header. It also repeats the plain leg, which is how a false positive got caught before it became a sentence: two servers appeared to change behaviour (409 against an aborted first leg), and the repeat plain leg answered the same 409. The first version of that instrument had no repeat leg and classified both as a header effect - a difference between two requests sent seconds apart, mistaken for a difference caused by one header. Ten servers rejected the GET identically either way; three answered GET with application/json and no stream at all, again identically. The write-up, including the false positive the repeat leg caught, is Can a header redirect a server to a method the body never asked for?

8. Does a server ever say how long its tool list may be cached?

A gateway has to decide how long to keep a cached list, and the spec lets the server answer that with ttlMs and cacheScope - but only from protocol revision 2026-07-28 onward, which decides what a count of silent servers can mean. Asked at 2025-06-18, where those fields are not required, 17 of 40 endpoints returned a list and 16 declared neither, 1 declared both (ttlMs: 300000 with cacheScope: "private"). Re-asked the same day, over the exactly identical 40 endpoints, at 2026-07-28: 13 returned a list, only 3 accepted that revision, and 0 of those 3 sent the fields it requires. So the claim that servers do not declare hints survives only in the small place where it can actually be tested - and the larger count, read as conformance, was never evidence of anything. We were reading neither field ourselves, caching every upstream for a hard-coded 60 seconds in a process-global map shared by all tenants; measured through a whole-frame stdio tee once the new revision is negotiated, our server does send both.

Joining the two same-day legs by endpoint explains why the denominator collapsed: 5 servers answer a 2025-06-18 initialize - several with a full tool list - and return HTTP 400 to a 2026-07-28 one instead of replying with their latest supported revision. A 400 gives a client nothing to downgrade from, so those endpoints are not legacy-era to a newer client, they are simply absent. That is why only 3 of the 40 could be asked about the new fields at all, and it is a shape worth knowing before any client ships a validator that treats absence as invalid.

The timing half is small: one declarant in sixteen, and honouring it meant clamping to 1 s..10 min so that a declared 0 cannot turn a read into a fresh handshake and a declared decade cannot freeze a list. The sharing half is the one worth the change, and it is stated without inflation: listTools() takes no caller identity, so nothing was exposed - the cache was merely built as though sharing were always safe, which is the shape that would turn a future per-caller header into a cross-tenant leak with nobody objecting. The write-up, its blind spots, and our own silence.

Three more things the same API told us

Reading registry records for their own reason surfaced a separate question: does a listing carry an artifact an installer can consume? On the alphabetically-first 54 records the answer was 6 of 54 carried a package at all, and that page records why counting them on the oldest list row instead of the latest record gives a different number. Walking the whole default list rather than its front half - 37,013 servers visible, 37,854 when removed records are asked for - puts that share at 41.85%, and finds 439 records (1.20%) that declare neither a package nor a hosted endpoint, which is the only group a client genuinely cannot act on. A package and a remote are two different ways to be actionable, and the first version of the sample page conflated them; the census page records the retraction and the reading that disproved it. Then a seeded draw of 200 of the 9,896 npm-listed records came back the other way: 196 resolve on npm at exactly the listed version, 3 list a version npm does not have, and 1 name is gone - a 2.00% unusable rate with a 95% interval of 0.06% to 3.94%, which is the number that narrows the problem to the records with no coordinates at all rather than stale ones. The other five artifact types were probed the same way afterwards - 15 of 785 do not resolve, with cargo and NuGet counted in full - and pooled across all six types the unusable rate is 1.93% (Wilson 1.24% to 2.99%). Scoring our own server with the rubric a directory uses to score strangers is a separate reading, and the honest version of it is on this page: 15 of 15 of our tools declare no outputSchema, while all 15 do carry titles and all four MCP annotations.

The numbers as a file

Every table on this page also exists machine-readable: docs/data/mcp-ecosystem-measurements.json - one entry per question, each with its script, its verdict counts and the per-endpoint rows behind them. Regenerate it with

node tools/build-mcp-measurement-dataset.mjs 40 docs/data/mcp-ecosystem-measurements.json

The file is generated and never hand-edited, and the builder refuses to write it if any question's verdict tally does not sum to that question's own row count. That guard is here because prose on this site has already disagreed with the table beneath it once (#180): a headline said 2/19 next to a tally of 18, and every doc check in the repository passed, because they match words and none of them does arithmetic. A generated file cannot make that particular mistake quietly.

It contains verdicts, counts, character lengths and a SHA-256 prefix for each instructions value. It contains no server response text: an instructions field is precisely the kind of payload that should not be copied into an artifact and re-distributed, and a dataset about that problem is not going to demonstrate it.

Two files, two samples, one command. mcp-ecosystem-measurements.json is the 40-endpoint run the tables above are drawn from (19:04-19:08 UTC). mcp-ecosystem-measurements.wide.json is the widest slice these scripts can reach - 54 to 55 unique endpoints out of the first 400 registry records, 19:15-19:22 UTC - and it is the run behind the figures quoted in two upstream threads on 2026-09-27: 15 of 23 servers send instructions (72 to 19,687 characters, median 406, one of fifteen over 2,000) and 1 of 23 answers server/discover while 19 return -32601. Both are produced by the same command with a different limit.

The row counts in the wide file differ between questions (54 or 55), because each of the five scripts re-reads the registry on its own and nothing pins the list across the five runs; the cause of the one-endpoint difference is not recorded. It is left visible in the files rather than smoothed over, since a dataset that presents its own sample as fixed is one nobody can re-run honestly.

Method, in one paragraph

Each script reads the official MCP registry for streamable-http records, takes the first 40 endpoints in the registry's default order, and sends anonymous JSON-RPC: initialize, then notifications/initialized, then tools/list - or, for the fourth script, server/discover - with a second page or a header-omitted repeat only where the specific question requires it. No credentials, no writes, no tool invocation, sequential requests with a short delay. The two final requests in the session probe differ in exactly one header, so a rejection cannot be blamed on something else the client forgot to send.

node tools/survey-mcp-tools-list-pagination.mjs 40
node tools/survey-mcp-revision-tolerance.mjs 40
node tools/survey-mcp-session-enforcement.mjs 40
node tools/survey-mcp-server-discover.mjs 40
node tools/survey-mcp-instructions-field.mjs 40
node tools/survey-mcp-protocol-version-header.mjs 40
node tools/survey-mcp-route-headers.mjs 40
node tools/survey-mcp-get-stream-headers.mjs 40
node tools/survey-mcp-list-cache-hints.mjs 40

We ran the same three questions against ourselves

Publishing numbers about other people's servers is easy; the test of whether the questions mean anything is answering them about your own. Our HTTP MCP endpoint (pnpm mcp:http, loopback only, credential-free default, port 3299 for the run) was asked the same three things on 2026-09-27 17:19 UTC:

Our own server, same probes
QuestionOur answerWhich group that puts us in
Asked for 2025-06-18 negotiated 2025-06-18, HTTP 200 normal
Does our tools/list paginate? 15 tools, no nextCursor we are in the non-paginating majority ourselves - which is the honest context for the 0 of 16 above
Do we issue a session id? none issued; requests served without one the 14-of-16 stateless group
Asked for 9999-99-99 answered 2025-11-25, HTTP 200 we substitute a supported revision rather than echo the impossible one - we are not in the group of 2

The fourth row is also the clearest illustration of issue #178: our server's fallback revision (2025-11-25) is newer than the one our own gateway client declares to upstreams (2025-06-18), and when these probes ran the gateway did not read what came back. Two components of the same product disagreed about which protocol they were on, and nothing in the code noticed. Both halves of that are now different in master and in the rolling :latest / :master image tags built from it on 2026-09-27, though not in the versioned 0.8.0 tag: the governed upstream client records the revision each server answered and names it per server in the servers array of GET /mcp/tools, and the second MCP client path in claude-code-patterns sends the same constant as the first, with a test that fails if the two drift apart. What the gateway should do when a server names a different revision - refuse it, or serve it and say so - is still open on the issue, and this page will be edited again when it is decided.

The run left nothing behind: the port was probed again after shutdown and refuses connections, and no http-entry process remained.

Limits, stated rather than implied