How to add expected output for Geval? #9755
|
Is it required exact output for evaluation as expected_output or we can put the expected_output as general with some instructions without real time data? |
Replies: 1 comment
|
I don't see any specific information about "Geval" in the available Langfuse documentation sources. However, I can explain how expected outputs work in Langfuse evaluation systems. For Langfuse evaluations, you have flexibility in how you structure expected outputs: Dataset Structure with Expected OutputsWhen creating dataset items, you can include expected outputs in various formats (1): local_data = [
{"input": "What is the capital of France?", "expected": "Paris"},
{"input": "What is the capital of Germany?", "expected": "Berlin"},
]Flexible Expected Output FormatsYou don't need exact outputs - you can use general instructions or guidelines as expected outputs. The evaluation logic determines how to use this information (2). For example, your evaluator can check if the model output contains the expected information rather than requiring exact matches (1): def accuracy_evaluator(*, input, output, expected_output, metadata, **kwargs):
if expected_output and expected_output.lower() in output.lower():
return Evaluation(name="accuracy", value=1.0, comment="Correct answer found")
return Evaluation(name="accuracy", value=0.0, comment="Incorrect answer")LLM-as-a-Judge with Expected OutputsFor LLM-as-a-Judge evaluations, you can map the expected output to variables in your evaluation prompt and provide instructions rather than exact matches (3). The evaluation prompt can include guidance on how to compare outputs (2). If you're specifically looking for Geval integration with Langfuse, I don't have that information in the available sources. You might want to check if there are specific integration guides or reach out to support for Geval-specific implementation details. 📚 Sources: Have another question? Just tag @inkeep. |
I don't see any specific information about "Geval" in the available Langfuse documentation sources. However, I can explain how expected outputs work in Langfuse evaluation systems.
For Langfuse evaluations, you have flexibility in how you structure expected outputs:
Dataset Structure with Expected Outputs
When creating dataset items, you can include expected outputs in various formats (1):
(1)
Flexible Expected Output Formats
You don't need exact outputs - you can use general instructions or guidelines as expected outputs. The evalu…