Repository navigation
Test rounds
Before and during the beta we ask the production server large batches of real questions, the way people ask an assistant, and read every answer. These pages publish what those rounds found, including what went wrong.
| Round | Server | Questions | How it was asked | What it judged |
|---|---|---|---|---|
| Test round 2026-09-21 v0.1.1 | 0.1.1 | 1266 | Each question mapped by hand to a tool call; no model in the loop | Whether each answer was what the question needed, plus latency |
| Test round 2026-10-04 v1.7.0 | 1.7.0 | 3000 | Four blind assistants picked tools from the catalogue alone; then robustness, protocol and load probes | Routing, answer structure and plausibility, robustness, performance |
| Test round 2026-10-06 v1.8.17 | 1.8.17 | 1000 | Ten blind assistants asked, German and English, in random order | Right tool, right arguments, a correct and complete answer, credit line, language |
| Test round 2026-10-07 v1.8.29 | 1.8.29 | 1000 | The same 1000 questions again, same order, same answer key | The same measures, so the two rounds compare directly |
The last two rounds asked the same 1000 questions and graded them the same way, so their numbers compare directly:
| Measure | 1.8.17 | 1.8.29 |
|---|---|---|
| Right tool | 99.1 % | 99.6 % |
| Right arguments | 98.1 % | 98.6 % |
| Correct and complete answer | 73.9 % | 91.6 % |
| Credit line on every answer with data | 99.9 % | 100 % |
| Answer in the asked language | 94.0 % | 99.8 % |
| Major findings | 200 | 71 |
| Critical findings | 12 | 0 |
| Latency per call, median / p95 | 394 / 4285 ms | 455 / 3255 ms |
The two earlier rounds measured different things, so only some of their numbers line up:
| Measure | 0.1.1 | 1.7.0 | 1.8.17 | 1.8.29 |
|---|---|---|---|---|
| Right tool, picked blind from the catalogue | – | 99.9 % | 99.1 % | 99.6 % |
| Credit line present | 94.0 % | 99.2 % | 99.9 % | 100 % |
| Place names that came back "which one do you mean?" | 6.9 % | 4.9 % | – | – |
| Server errors (5xx) or transport errors | 0 | 0 | 0 | 0 |
| Latency per call, median | 149 ms | 332 ms | 394 ms | 455 ms |
Answers have become richer since 0.1.1 (coordinates, the Land, ids for the next question, labelled provider text), and the median latency has grown with them. It is still under half a second.
-
Production, from outside. Every round asked the public endpoint
https://mcp.viafrei.de/mcpover the internet, from one machine: the 1.7.0 round in the afternoon, the others in the evening or at night. - Blind. In the last three rounds the assistant choosing the tool saw only the question and the public tool catalogue, never the answer key.
- Graded by assistants. The 1.8.x rounds were graded by assistants against a written key. The askers also counted their own mistakes, so a few findings are about the asking, not the server.
- Text only. The graders read the answer text. A detail that exists only in the structured result does not count.
- Nobody checked the providers. No price, jam or departure was compared with its provider in real time. Grading judged plausibility: the right place, the right kind of data, a sensible age.
Something wrong in an answer you got? Report it; that is how the next round's questions get written.
For travellers
- Use case Driving Munich to Berlin
- Use case The commute that broke
- Use case An EV on a long weekend
- Use case The lorry driver and the legal break
- Use case First road trip in Germany
For developers
- Use case Building a local guide agent
- Use case A commuter bot that speaks up
- Use case Embedding ViaFrei in a travel app
For businesses
- Use case Fleet and logistics briefings
- Use case Hotels and tourist information desks
- Use case Getting thousands to an event
- Test rounds
- Test round 2026-10-07 v1.8.29
- Test round 2026-10-06 v1.8.17
- Test round 2026-10-04 v1.7.0
- Test round 2026-09-21 v0.1.1
https://mcp.viafrei.de/mcp
https://mcp.viafrei.de/sse
No key. No account.