SaSameKnowledge
publishedresearchresearch

How should MCP measurements be interpreted?

This page defines the language used for reachability, compatibility, claims and longitudinal changes.

Collection
research
Updated
2026-07-29
min read
2
Version
v1
Verified summary

Measurements are bounded results tied to an instrument, target identity, time and denominator; they are not timeless ratings.

Evidence rigor
Observed
2026-07-23 14:08 UTC
Claim strength
measured
Dataset scope
Registry-Runtime Gap headline measurements: 152,781 repeated observations, 4,162 enumerable tool surfaces, 78,041 declared tools, 78.2% overall annotation coverage, 30,641 potentially state-changing tools, of which 7,618 lack measured annotations.
Methodology
Each figure is defined by an explicit formula, source record set and exclusion rule in the published methodology, and is re-derived by an independent reproduction script rather than quoted from the paper text alone.
Limitations
Repeated observations, tool-surface counts, declared-tool counts and the coverage percentage reproduce exactly on re-run. The state-changing-tools-without-annotations figure does not currently reproduce, because the original classifier's exact provenance was not retained (NOT_REPRODUCIBLE).

Measured — backed by a specific dated observation or dataset

01

Purpose

This page defines the language used for reachability, compatibility, claims and longitudinal changes.

02

Method

Each metric documents formula, source records, exclusions and freshness state.

MetricValueReproduces?
Repeated observations152,781yes
Enumerable tool surfaces4,162yes
Declared tools78,041yes
Overall annotation coverage78.2%yes
Potentially state-changing tools30,641not verified
State-changing tools without annotations7,618no (classifier provenance lost)
03

Boundary and limitations

A result can be valid for its sample while being inappropriate for a different population.

EX

Examples

2
  1. 01

    Runtime availability among observed endpoints is not reported as the percentage of all MCP servers in existence.

  2. 02

    78.2% overall annotation coverage means roughly one in five declared tools in the sample carried no measured runtime annotation at observation time — the measurement is scoped to the sample, not extrapolated to "the MCP ecosystem" as a whole.

RF

References

1