feat(experiments): Support splitting task JSON response into output and metadata fields in dataset runs #11214
Unanswered
anuradhapalakurthiTR
asked this question in
Support
Replies: 1 comment 1 reply
|
Hi @anuradhapalakurthiTR, thanks! You can use the Langfuse SDK within your task function to control and manually set the output of your observation or trace. I have created an example here: from langfuse import get_client
from langfuse.openai import OpenAI
# Initialise client
langfuse = get_client()
# Define your task function
def my_task(*, item, **kwargs):
with langfuse.start_as_current_observation(as_type="span", name="my-span") as span:
question = item["input"]
response = OpenAI().chat.completions.create(
model="gpt-4.1", messages=[{"role": "user", "content": question}]
)
# Set the output and metadata on the observation level
span.update(
output=response.choices[0].message.content,
metadata=response
)
# Set the output and metadata on the observation level
span.update_trace(
output=response.choices[0].message.content,
metadata=response
)
return response.choices[0].message.content
# Run experiment on local data
local_data = [
{"input": "What is the capital of France?", "expected_output": "Paris"},
{"input": "What is the capital of Germany?", "expected_output": "Berlin"},
]
result = langfuse.run_experiment(
name="Geography Quiz",
description="Testing basic functionality",
data=local_data,
task=my_task,
)
# Use format method to display results
print(result.format())Let me know if this works for you! |
1 reply
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Describe your question
Feature Request
Use Case:
When running experiments via the Experiment Runner SDK, my task function returns a JSON response with multiple fields. I want to separate this response so that:
message) appears as the Output in the dataset runCurrent Behavior:
The entire JSON response from the task function is stored as the output in dataset runs. There's no documented way to automatically split the response into output vs metadata fields.
Desired Behavior:
Provide a way to specify which fields from a task's JSON response should be treated as output vs metadata, either through:
run_experiment()parametersExample:
All reactions