Skip to content

Test rounds

Sergii Mavrov edited this page Oct 7, 2026 · 1 revision

Test rounds

Before and during the beta we ask the production server large batches of real questions, the way people ask an assistant, and read every answer. These pages publish what those rounds found, including what went wrong.

Round Server Questions How it was asked What it judged
Test round 2026-09-21 v0.1.1 0.1.1 1266 Each question mapped by hand to a tool call; no model in the loop Whether each answer was what the question needed, plus latency
Test round 2026-10-04 v1.7.0 1.7.0 3000 Four blind assistants picked tools from the catalogue alone; then robustness, protocol and load probes Routing, answer structure and plausibility, robustness, performance
Test round 2026-10-06 v1.8.17 1.8.17 1000 Ten blind assistants asked, German and English, in random order Right tool, right arguments, a correct and complete answer, credit line, language
Test round 2026-10-07 v1.8.29 1.8.29 1000 The same 1000 questions again, same order, same answer key The same measures, so the two rounds compare directly

The trend

The last two rounds asked the same 1000 questions and graded them the same way, so their numbers compare directly:

Measure 1.8.17 1.8.29
Right tool 99.1 % 99.6 %
Right arguments 98.1 % 98.6 %
Correct and complete answer 73.9 % 91.6 %
Credit line on every answer with data 99.9 % 100 %
Answer in the asked language 94.0 % 99.8 %
Major findings 200 71
Critical findings 12 0
Latency per call, median / p95 394 / 4285 ms 455 / 3255 ms

The two earlier rounds measured different things, so only some of their numbers line up:

Measure 0.1.1 1.7.0 1.8.17 1.8.29
Right tool, picked blind from the catalogue – 99.9 % 99.1 % 99.6 %
Credit line present 94.0 % 99.2 % 99.9 % 100 %
Place names that came back "which one do you mean?" 6.9 % 4.9 % – –
Server errors (5xx) or transport errors 0 0 0 0
Latency per call, median 149 ms 332 ms 394 ms 455 ms

Answers have become richer since 0.1.1 (coordinates, the Land, ids for the next question, labelled provider text), and the median latency has grown with them. It is still under half a second.

How to read these pages

  • Production, from outside. Every round asked the public endpoint https://mcp.viafrei.de/mcp over the internet, from one machine: the 1.7.0 round in the afternoon, the others in the evening or at night.
  • Blind. In the last three rounds the assistant choosing the tool saw only the question and the public tool catalogue, never the answer key.
  • Graded by assistants. The 1.8.x rounds were graded by assistants against a written key. The askers also counted their own mistakes, so a few findings are about the asking, not the server.
  • Text only. The graders read the answer text. A detail that exists only in the structured result does not count.
  • Nobody checked the providers. No price, jam or departure was compared with its provider in real time. Grading judged plausibility: the right place, the right kind of data, a sensible age.

Something wrong in an answer you got? Report it; that is how the next round's questions get written.

Clone this wiki locally