Replies: 2 comments
|
There is no explicit support documented for configuring dataset.run_experiment() evaluators to automatically access and evaluate nested spans. Evaluators in run_experiment() receive the input, output, expected output, and metadata for each dataset item, which typically relates to the top-level trace or main task output(1). The examples and API reference for evaluators focus on evaluating the outputs at the item level and do not show mechanisms for traversing or programmatically targeting nested spans(1). Your current approach—calling span.score() within the retrieval span—remains the only documented way to attach evaluations directly to nested spans(1). There is no indication in the sources that run_experiment() evaluators can be configured to reach into nested spans automatically, nor is there an official API for enumerating or passing all nested observations to evaluators within the experiment runner(1). If you want to keep evaluation logic out of your main execution code, you would need to continue using manual scoring or build custom logic outside the experiment runner to traverse and evaluate nested spans after trace creation(1). 📚 Sources: Have another question? Just tag @inkeep. |
|
Hi @Xixaus, currently, the evaluation method of the experiment runner SDK can only access trace-level information. One option would be to add the information you want to evaluate to the metadata, which can be accessed by the eval methods. Alternatively, you would need to use .score() if you want to evaluate nested spans. |
Uh oh!
There was an error while loading. Please reload this page.
Hi! I'm using dataset.run_experiment() with evaluators, but they only seem to receive data from the top-level trace.
I need to evaluate nested observations - specifically a document retrieval step. Currently, I'm calling span.score() directly within the retrieval span, but this hardcodes evaluation logic into my execution code.
Is there a way to configure evaluators in run_experiment() to access and evaluate nested spans? Or do I need to keep using the manual span.score() approach?
Thanks!
All reactions