Skip to content
Discussion options

You must be logged in to vote

Unfortunately, the Langfuse-hosted LLM-as-a-Judge evaluators don't currently support iterating over a list of questions within a single evaluator to compute an average score(1).

The current architecture processes each dataset item as a single evaluation pass—the evaluator prompt is populated with your mapped variables (like {{llm_result}} and {{questions}}), sent to the LLM once, and expects a single score and reasoning in return(1).

Your options:

  1. External evaluator (client-side): As you mentioned, you can implement custom scoring logic via the SDK that iterates through questions, calls the judge for each, averages the results, and then submits the final score back to Langfuse. See the …

Replies: 6 comments 3 replies

Comment options

You must be logged in to vote
0 replies
Answer selected by lleviraz
Comment options

You must be logged in to vote
0 replies
Comment options

You must be logged in to vote
0 replies
Comment options

You must be logged in to vote
2 replies
@nikhildigde
Comment options

@jannikmaierhoefer
Comment options

Comment options

You must be logged in to vote
1 reply
@jannikmaierhoefer
Comment options

Comment options

You must be logged in to vote
0 replies
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
3 participants