What should TraceArena measure in your agent world? | 你的 Agent 世界最需要验证什么? #3
Replies: 1 comment
-
|
v0.1.6 adds one concrete lesson from the first replay report: a run needs two different hashes. That distinction makes the evaluation question more precise: should a scenario optimize for semantic determinism, recovery quality, evidence completeness, or all three? If you have an agent workflow, please share the outcome that must be trusted and the failure that would be most expensive to miss. Release: https://github.com/tonyhyworld/TraceArena/releases/tag/v0.1.6 |
Beta Was this translation helpful? Give feedback.
Uh oh!
There was an error while loading. Please reload this page.
-
TraceArena is now public. The first no-key Market Replay is deliberately small: it lets you inspect the runtime path from evidence and structured actions to world events, settlement, and replay without an API key or brokerage connection.
We are using this thread to shape the next public scenarios and evaluation adapters.
What would you most like to test?
If you have an existing Agent, please include:
Links:
Please share constraints and failures as well as successful demos. We will summarize recurring requests into public issues and scenario-pack priorities.
Beta Was this translation helpful? Give feedback.
All reactions