Dataset run items showing identical scores despite unique observation-level scores #11513
Replies: 4 comments 8 replies
|
Based on your description, you're creating observation-level scores correctly and linking them to dataset run items via The Langfuse data model supports linking scores to observations within traces(2), and when you create a Potential causes:
Recommendations:
langfuse.create_score(
trace_id=field_span.trace_id,
observation_id=field_span.id,
name=metric_name,
value=metric_value,
data_type="NUMERIC",
)
# Fetch scores for specific observations
scores = langfuse.api.scores.get(
trace_id="your_trace_id",
observation_id="your_observation_id"
)If observation-level scores are correct but dataset run items still show identical values, this may indicate a display or aggregation issue in the UI for your version (v3.146.0). Consider checking if there are any updates or known issues for your specific version, or reach out to Langfuse support with your specific version details. 📚 Sources:
Have another question? Just tag @inkeep. |
|
Hi @mscherrmann, what is the reason you are creating dataset run items manually and not using the experiment runner sdk or the low-level method? Manually creating dataset run items is not recommended, as this might cause problems, as a score should be linked both to the observation and the dataset run item. |
|
We are running into the same issue. Our issue is that, we uses Nestjs and BullMQ, and we put the evaluation task into a bullmq that's instrumented with opentelemetry. Using experiment runner sdk will group all the spans under one root trace created by bullmq. Because of this, all our run items show the identical score (which is the average score of all spans of a single root). The cost is wrong too, each experiment item receives the total of cost. Any solution to this issue? |


Uh oh!
There was an error while loading. Please reload this page.
Describe your question
I'm experiencing an issue where dataset run items display identical scores across all items, even though the underlying observations have correct, unique score values.
Setup
I'm creating dataset runs by linking observations to dataset items using the API:
Before creating the dataset run, I push observation-level scores:
I call langfuse.flush() after pushing all scores and before creating dataset run items.
Dataset run view:

Dataset item run view (individual scores equal to dataset average scores):

Langfuse Cloud or Self-Hosted?
Self-Hosted
If Self-Hosted
v3.146.0
If Langfuse Cloud
No response
SDK and integration versions
langfuse-python-sdk version: 3.11.2
Pre-Submission Checklist
All reactions