Replies: 3 comments 3 replies
|
Yes — that matches our run, and I'd add three fields rather than drop any. For the record, ours in your frame: sampling frame and draw — 50 tenants per platform from the exact-label rows of a public technology-detection dataset (August snapshot), one seed (20260917), one UTC date, 7,774 requests. Paths, in order — robots.txt, platform fingerprint, the well-known sweep (/.well-known/ucp, /.well-known/mcp/server-card.json, /.well-known/agent-card.json, /.well-known/api-catalog, …), then JSON-RPC handshakes on any endpoint discovered. Positive — an MCP endpoint that answers tools/list; a parseable profile alone is not positive. Exclusions, counted — 4 robots refusals (Disallow: /); 21 "could not ask", of which 18 were our own concurrency colliding with one platform's edge and were re-asked serially (all 18 genuine negatives), the remaining 3 unreachable. Every exclusion is reported beside the count, never folded into "no". UCP version per positive and both platforms' full tool signatures are in our METHOD.md. Server card — recorded: all 96 positives were found through /.well-known/ucp; none through the server card. That last point is the reason for my first added field. Three fields I'd add:
One small disagreement: I'd keep the positive criterion as "protocol response" for the MCP surface and treat "valid profile document" as a separate, weaker row rather than an alternative — the same reasoning as your HTML-200 point. Happy to review the draft. We're running a larger pass now that records every field above per host; when it's done I'll publish the frame fields alongside the counts so the two can be compared directly. |
|
Thanks Francesco — this is the frame, and I'd use it as written. One addition to "could not ask": split it by whose side the reason is on. robots refusals, dead DNS, TLS failures and 5xx are the site's; 429 with a Retry-After, and timeouts on an origin we were sharing with other hosts in the same run, are plausibly the prober's. We report the second group as its own line because it is the only one the prober can reduce, and folding the two together hides whether a run was polite. Two small things on the rows. Row 4 should carry the discovery route per positive, as your run-fields table already asks, so that "answered tools/list via the profile's endpoint" and "answered via a server card" stay separable. And for row 3, naming the schema commit is what made your 15/15 re-check possible — keep that mandatory. Your note that the earlier "valid" row was a JSON-parse check is exactly the kind of thing this frame makes visible, and re-running it is the right response. We'll fill the frame in for our next published pass rather than quote numbers that haven't been through our own checks yet. Dean |
|
Updated reporting frame, folding in Dean's review. Run fields
Both "could not ask" lines are reported beside the result and never folded into "no". Outcome rows (hosts, not requests)
A row is only filled if the run measured it. "Not measured" is a valid entry, a guess is not. Our own run, filled in against the first draft, is here; it does not split "could not ask" by side yet. Francesco Marinoni Moretto |
Uh oh!
There was an error while loading. Please reload this page.
Three independent measurements of the live UCP/MCP surface have been published in the last month, and a fourth has been announced:
tools/listtools/list/.well-known/ucpdocument (GET-only, no MCP calls)The results agree on the direction, but the numbers cannot go in one table: 49 of 70 and 12 of 50 are not the same measurement pointing at different values, they are different measurements. They differ on at least three things:
tools/listthat answers, others a UCP profile that parses. A200alone is not enough: in our run, 32 responses with status 200 on agentic paths were HTML error pages. Counted as positives, they would have overstated adoption by roughly a third.robots.txtDisallow: /), and three of six controls went unmeasured (one 429, two off-host redirects the instrument does not follow). Unless those hosts are reported, two identical populations can give different rates.Proposal: a short, shared reporting frame, not a shared tool. Each census keeps its own method and states, in the same fields:
versioneach positive declares;keysvssigning_keys, per @westonale's point in UCP Readiness Checker - an open, API-first validator that grades any store A–F #656), and whether/.well-known/mcp/server-card.jsonreturned a valid document.With that, the next round is comparable, and a disagreement between two censuses means something instead of reflecting two frames.
I'm happy to write the first draft as a markdown file others can edit or reject.
Before that, it would help to hear from @krisdiallo, @deangoodman and @unblinkr: does this match how you would describe your own run, and is there a field you would drop or add?
All reactions