Langfuse Eval workflow #11432
Describe your questionI have a questions and expected outputs pairs file. I would like to run eval comparing the expected outputs and trace outputs and giving a score. I'm circling and cannot find out how I can do it. Langfuse Cloud or Self-Hosted?Self-Hosted If Self-Hosted3.142.0 If Langfuse CloudNo response SDK and integration versionsNo response Pre-Submission Checklist
|
Replies: 2 comments 2 replies
|
Based on your question-expected output pairs file, you can use Langfuse's Experiments via SDK to evaluate traces against expected outputs.(1) Recommended Approach: Experiments via SDKThe SDK provides an experiment runner that automatically handles evaluation. Here's how to set it up:(1) Python SDK Example: from langfuse import get_client, Evaluation
# Initialize client
langfuse = get_client()
# Define your task function (this runs your application)
def my_task(*, item, **kwargs):
question = item["input"]
# Your application logic here that produces output
# ...
return output
# Define evaluator to compare expected vs actual output
def accuracy_evaluator(*, input, output, expected_output, metadata, **kwargs):
if expected_output and expected_output.lower() in output.lower():
return Evaluation(name="accuracy", value=1.0, comment="Correct answer found")
return Evaluation(name="accuracy", value=0.0, comment="Incorrect answer")
# Load your questions and expected outputs
test_data = [
{"input": "What is the capital of France?", "expected_output": "Paris"},
{"input": "What is the capital of Germany?", "expected_output": "Berlin"},
]
# Run experiment with evaluation
result = langfuse.run_experiment(
name="My Evaluation",
data=test_data,
task=my_task,
evaluators=[accuracy_evaluator]
)
print(result.format())Key Points:
The evaluator function returns an 📚 Sources: Have another question? Just tag @inkeep. |
|
Hi @minkhantkyaw-brillar, you can follow the instructions from Inkeep bot above, or run an experiment via the Langfuse UI:
|
Based on your question-expected output pairs file, you can use Langfuse's Experiments via SDK to evaluate traces against expected outputs.(1)
Recommended Approach: Experiments via SDK
The SDK provides an experiment runner that automatically handles evaluation. Here's how to set it up:(1)
Python SDK Example: