TIMING NOTE / SEPTEMBER 2026
Why an MCP client gives up before the server answers
A stdio MCP server has to be a running process before it can say anything, and a client's connect budget starts when the client spawns it, not when the server finishes loading. When those two facts meet, the failure looks like a broken server: the client reports a timeout, the server never gets a chance to answer, and nothing in either log says "I was still importing". We measured our own server against that budget so the number is on the record instead of in a guess.
The numbers
Run on 2026-09-27 on a Windows 10 x64 machine with Node.js v25.8.1, from a clone of the
repository, against the source entry point the published image also runs
(packages/mcp-server/src/index.js):
node tools/mcp-startup-profile.mjs --repeat 5
| Run | initialize |
tools/list |
Gap | Tools |
|---|---|---|---|---|
| 1 | 8,461 ms | 8,464 ms | 3 ms | 15 |
| 2 | 7,853 ms | 7,860 ms | 7 ms | 15 |
| 3 | 8,304 ms | 8,309 ms | 5 ms | 15 |
| 4 | 7,630 ms | 7,633 ms | 3 ms | 15 |
| 5 | 7,787 ms | 7,790 ms | 3 ms | 15 |
Five consecutive attempts answered initialize between
7,630 ms and 8,461 ms (median 7,853 ms, spread 831 ms). In every run the
complete 15-name tool list arrived 3 to 7 ms later. The server identified
itself as unified-ai-system over protocol 2025-06-18, and wrote
69 bytes to stderr, identically in all five runs.
A second and third batch the same evening, same command, read lower.
The table below is verbatim from two later runs of
node tools/mcp-startup-profile.mjs --repeat N --json against the same entry
with the same stripped environment:
| batch | time (UTC) | runs | initialize range |
median | spread |
|---|---|---|---|---|---|
| A (above) | 2026-09-27 ~17:19 | 5 | 7,630 – 8,461 ms | 7,853 ms | 831 ms |
| B | 2026-09-27 21:37 | 3 | 6,399 – 6,671 ms | 6,429 ms | 272 ms |
| C | 2026-09-27 21:40 | 5 | 6,391 – 6,688 ms | 6,654 ms | 297 ms |
B and C agree with each other and both sit entirely below A's minimum: a gap of roughly
1.2–1.8 s, which is about four times the internal spread of either batch. So this is a
difference between batches, not noise within one - and a real difference
is exactly the kind of thing that gets quietly written up as an improvement.
This page does not claim startup got faster. When this section was first
published, two explanations were live and indistinguishable from these numbers alone:
load at the time A was taken (batch C ran on a machine with nine node.exe
processes alive, so it was not measured on an idle box either), or the ESM cycle cuts that
landed in this repository the same day, which shorten the import chain that has to resolve
before the first byte and so could plausibly be worth a second and a half.
The code-change leg has now been run, and it does not explain the gap.
Twenty-six of the thirty-five paths the cut touched were restored from its parent commit
in this worktree - path-wise and in place, not a copied or junctioned tree, because a moved
pnpm tree is known on this machine to produce false timeouts - and five cold
boots of each state were timed back to back with the same stripped environment on
2026-09-28 (~00:25 UTC):
| leg | which code | runs | initialize range |
median | spread |
|---|---|---|---|---|---|
| D | with the cycle cut (HEAD) | 5 | 6,991 – 7,763 ms | 7,209 ms | 772 ms |
| E | pre-cut, same command | 5 | 6,948 – 8,914 ms | 7,252 ms | 1,966 ms |
The pre-cut leg was 43 ms slower, not 1.5 s faster - the wrong direction for the code hypothesis, so we drop it. Three limits belong in the same breath: the legs ran consecutively rather than interleaved; the third leg, a repeat of D after restoring the tree, aborted on a pathspec error before it measured anything, so this is one pair and not the replicated A/B/A that was sketched; and E's own internal spread, 1,966 ms, is larger than the gap it was asked to explain. Five boots per leg could not have resolved a 1.5 s difference even if one were there. Read this as the code explanation is not supported, not as the effect is proven absent.
What the pair does change is the number this page publishes. The pre-cut tree, on a quieter
machine than batch A, still read 7,252 ms - above batch C's 6,654 - so cold start here
wanders by more than a second with no code change at all. The range therefore widens rather
than tightens: across the 23 timed boots in the two tables above, 6.4–8.9 s
to answer initialize, and every one of those boots still returned the same 15
tool names. On a laptop, read the wider range, not the prettier one.
One more experiment moved this from "slow" to "fragile", and it is the reason the number above should not be read as a latency complaint. Running the same three-repeat profile while eight CPU burners were busy: 0 of 3 runs answered at all within 45,000 ms - not slower, absent. Kill the burners and measure again immediately: 3 of 3 answered at 6,618-6,669 ms. A loaded machine does not make this handshake 20% worse; past some point it makes it not happen, and the effect reverses as soon as the load leaves. Three runs place no number on that point.
This is the shape of the real complaint, because the machine most likely to run this server is a laptop already compiling, indexing, or hosting other MCP servers. It is also why the 1.2-1.8 s gap between batches needs no code-change explanation to be plausible: load moves this metric by tens of seconds, so a gap the size of one second is well inside its reach. The pre-cut comparison has since been run - legs D and E above - and it did not favour the code explanation, so load is the account this page keeps; it still does not claim startup got faster.
Where the budget actually goes
The gap is the whole point. If enumerating tools were expensive, the second number would move. It does not: whatever is slow happens before the first response byte, which for a stdio server means process start plus the module graph, and not MCP work. Two consequences follow:
- A client whose connect budget is shorter than the import time cannot be helped by trimming the tool list, caching discovery, or making the server "lazy" about tools - the server is not the part that is slow yet.
- A server that answers in 7-8 seconds fits inside a 30-second budget but fails inside a 5-second one, so the same artifact is "fine" in one client and "failed" in the next. OpenCode documents a default MCP startup timeout of 30 s; the report that led us to build this measurement described a cold install that landed at 33-34 s and was marked failed on that boundary.
There is no first-launch premium visible here. Run 1 was the slowest of the five, so on this machine the seven seconds are paid on every launch, not once per cache lifetime - which is worth saying out loud, because "just warm the cache" is the advice most often given for this symptom.
Method, and how to run it against your own server
# our server, five attempts
node tools/mcp-startup-profile.mjs --repeat 5
# any stdio MCP server on your machine
node tools/mcp-startup-profile.mjs -- npx -y some-mcp-server@1.2.3
# machine-readable
node tools/mcp-startup-profile.mjs --repeat 3 --json
What the harness does, and deliberately does not do:
-
It spawns the command, sends
initialize, and sendstools/listthe instant the first answer is parsed, so the gap measures the server's own work rather than our write latency. -
It runs the child with only
PATHandNODE_ENV=productionin its environment, and refuses to start at all if the requested environment carries a variable whose name matchesKEY|TOKEN|SECRET|PASSWORD|CREDENTIAL|AUTH. The no-provider claim is enforced by the harness, not asserted by this page. -
It never treats a notification, a log line or a partial write as an answer: a line counts
only if it parses as JSON and carries an integer
id. A false "answered" is the one reading that would hide a real timeout. - It kills the child on every exit path, so a profiling run leaves no process behind.
What this page does not establish
- One machine, one Node version. Our CI pins Node 22 and this box ran v25.8.1, and we have separately watched a prebuilt-native dependency fail to provide binaries for a newer major - so do not read 6.4-8.9 s as a cross-platform constant.
- Not a comparison with other servers. We measured ours. Anybody can run the same command against their own and get their own five numbers; nothing here ranks anything.
- Not a claim that our startup time is acceptable. Seven to eight seconds to first byte is the reason this measurement exists: we would rather own the number in an issue (#168) than let a newcomer meet it as a "failed" badge in a client UI.