Skip to content
Discussion options

You must be logged in to vote

Based on your question-expected output pairs file, you can use Langfuse's Experiments via SDK to evaluate traces against expected outputs.(1)

Recommended Approach: Experiments via SDK

The SDK provides an experiment runner that automatically handles evaluation. Here's how to set it up:(1)

Python SDK Example:

from langfuse import get_client, Evaluation

# Initialize client
langfuse = get_client()

# Define your task function (this runs your application)
def my_task(*, item, **kwargs):
    question = item["input"]
    # Your application logic here that produces output
    # ...
    return output

# Define evaluator to compare expected vs actual output
def accuracy_evaluator(*, input, output, e…

Replies: 2 comments 2 replies

Comment options

You must be logged in to vote
0 replies
Answer selected by Lotte-Verheyden
Comment options

You must be logged in to vote
2 replies
@minkhantkyaw-brillar
Comment options

@Lotte-Verheyden
Comment options

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
2 participants