Dataset Experiment Run - Difficulty Running Judge Prompt for Multiple Questions in a Single Call #11499
Describe your questionHi, I’m running an experiment from the SDK on a dataset with X items, and for each item, and Evaluator is configured. ` Answering rules:
Model result: Questions: When there is a single question passed, it usually works ok, but if there are multiple questions, it becomes unreliable. The question is if there is a way to have the evaluator process the the questions one by one, and do some average, and only then report the score for that item (I believe it is possible to achieve if i run an external evaluator on the client side, but wanted to avoid that). Here is the answer from the great LF AI: https://langfuse.com/docs/ask-ai?conversationId=conv_01kes99zngf7ab9gn8m1p802ea Langfuse Cloud or Self-Hosted?Self-Hosted If Self-Hostedv3.66.1 OSS If Langfuse CloudNo response SDK and integration versionsNo response Pre-Submission Checklist
|
Replies: 6 comments 3 replies
|
Unfortunately, the Langfuse-hosted LLM-as-a-Judge evaluators don't currently support iterating over a list of questions within a single evaluator to compute an average score(1). The current architecture processes each dataset item as a single evaluation pass—the evaluator prompt is populated with your mapped variables (like Your options:
The feature you're looking for (iterative evaluation with aggregation within the Langfuse-hosted evaluator) isn't available in the current LLM-as-a-Judge implementation(1). If you need this functionality, the external evaluator approach on the client side is your best path forward. 📚 Sources: Have another question? Just tag @inkeep. |
|
Hi @lleviraz, thanks! Yes, inkeep is correct. This is not yet suppored. |
|
Ok, thanks.
…On Tue, Jan 13, 2026, 15:20 Jannik Maierhöfer ***@***.***> wrote:
Hi @lleviraz <https://github.com/lleviraz>, thanks! Yes, inkeep is
correct. This is not yet suppored.
—
Reply to this email directly, view it on GitHub
<#11499 (comment)>,
or unsubscribe
<https://github.com/notifications/unsubscribe-auth/AEQFPH4MQIWTYZUGGPSUHW34GTWLFAVCNFSM6AAAAACRNRVEUGVHI2DSMVQWIX3LMV43URDJONRXK43TNFXW4Q3PNVWWK3TUHMYTKNBYGU3DINQ>
.
You are receiving this because you were mentioned.Message ID:
***@***.***>
|
|
Sure, thank you. Can you please point us to the changes expected here ?
Nikhil Digde
…On Wed, 14 Jan, 2026, 3:43 pm Jannik Maierhöfer, ***@***.***> wrote:
Thanks @nikhildigde <https://github.com/nikhildigde>, we are currently
rebuilding the LLM-as-a-Judge logic in Langfuse, so it might make sense to
talk about a potential contribution once this is live in the next couple of
weeks.
—
Reply to this email directly, view it on GitHub
<#11499 (reply in thread)>,
or unsubscribe
<https://github.com/notifications/unsubscribe-auth/AJICMPCN56J4BYSP6YJM4LD4GYJE7AVCNFSM6AAAAACRNRVEUGVHI2DSMVQWIX3LMV43URDJONRXK43TNFXW4Q3PNVWWK3TUHMYTKNBZGQ2DAMA>
.
You are receiving this because you were mentioned.Message ID:
***@***.***>
|
|
Ok thanks
Nikhil Digde
…On Wed, 14 Jan, 2026, 4:14 pm Jannik Maierhöfer, ***@***.***> wrote:
@nikhildigde <https://github.com/nikhildigde>, we dont have any specs
written down publicly yet. But the high-level plan is to make the
evaluation decision based on the observation level and not the trace level
anymore.
—
Reply to this email directly, view it on GitHub
<#11499 (reply in thread)>,
or unsubscribe
<https://github.com/notifications/unsubscribe-auth/AJICMPEHUX2Y6RBHSGVUFYT4GYMYFAVCNFSM6AAAAACRNRVEUGVHI2DSMVQWIX3LMV43URDJONRXK43TNFXW4Q3PNVWWK3TUHMYTKNBZGQ3DKMA>
.
You are receiving this because you were mentioned.Message ID:
***@***.***>
|
Unfortunately, the Langfuse-hosted LLM-as-a-Judge evaluators don't currently support iterating over a list of questions within a single evaluator to compute an average score(1).
The current architecture processes each dataset item as a single evaluation pass—the evaluator prompt is populated with your mapped variables (like
{{llm_result}}and{{questions}}), sent to the LLM once, and expects a single score and reasoning in return(1).Your options:
External evaluator (client-side): As you mentioned, you can implement custom scoring logic via the SDK that iterates through questions, calls the judge for each, averages the results, and then submits the final score back to Langfuse. See the …