How do you measure the success of AX in a quantitative or qualitative manner? #5
Replies: 3 comments 4 replies
|
who’s evaluating AX—the user or the system owner? |
|
This is a super interesting question! I feel like we'll probably see some level of sort of super high-level grading system, similar to lighthouse or an SEO grader today, e.g. for things like:
And then at a higher sophistication level, maybe something like 'Agent Journey Tests' where you automate runs of the X most popular agents and evaluate % successful completion of core user workflows (e.g. what % of the time can the agent successfully make a purchase) and get aggregated results. Convex Evals is definitely good inspiration for this; the hard part, though, would be figuring out what the right tasks & tests are for each site (or each agent <> site combo if agents become more specialized) |
|
Overfitting and underfitting will be a critical part of this conversation as context from one system can be over/under, impacting the context to other parts of systems. For example, context can suggest using one provider over all others which obviously impacts other tools ability to play their role. |
Uh oh!
There was an error while loading. Please reload this page.
I have some ideas on this but I'm really curious what others are thinking about. Convex's evals is an inspiring start to this consideration
All reactions