Replies: 2 comments 2 replies
|
Based on your screenshots, you're experiencing duplicate evaluation scores - one on the trace level (1.00) and another on the span level (0.90) for the same dataset run item. When you configure an LLM-as-a-Judge evaluator with the target set to "New dataset run items"(1), the evaluator runs on traces created during dataset experiments(1). The evaluator configuration allows you to specify which dataset the evaluator should run on and whether to target new Dataset Runs(1). When using From your screenshots, it appears the evaluator is running twice - once evaluating the entire trace and once evaluating a specific span within that trace. This could happen if:
To troubleshoot this, you should check your evaluator's "Target filter" settings in the configuration(1). The sampling configuration targets "all future dataset run items that match these filters"(1). Additionally, there's a known issue where evaluator traces created via To resolve this, review your evaluator configuration to ensure it's only targeting the appropriate observation level (trace or span, not both). 📚 Sources:
Have another question? Just tag @inkeep. |



Uh oh!
There was an error while loading. Please reload this page.
Describe your question
I have a question for Datasets evaluattion. I set an Evaluator, Its target is New dataset run items. Then I dataset.run_experiment by SDK。 I checked the trace of dataset.Item.Run, There is an amazing happend which is both trace and span got a evaluation score, and they are different, Could you tell me what happend.Wish your response



Langfuse Cloud or Self-Hosted?
Langfuse Cloud
If Self-Hosted
No response
If Langfuse Cloud
No response
SDK and integration versions
No response
Pre-Submission Checklist
All reactions