Replies: 3 comments
|
With regard to the goal of item 3, consider two very simple tweaks to the current Human Annotation item view. There is a "Collapse all" button above each section. If the change made to those on one item persisted to the next item, I suspect most users would find that to be an improvement. If the "Expand all" version of that button expanded the JSON items inside the section, that seems like it would also be appreciated by most as well. In fact, it is probably the behavior most users expect. |
0 replies
|
@alexrosen thanks for the feedback here on our experimentation features. I will direct this to our engineers who work on experiments. We will look into this after the holidays. |
0 replies
|
Hey, thanks for sharing this extensive feedback!
|
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Describe the feature or potential improvement
I've been starting to do experiments on a prompt that has very subjective outputs. I want to use a Human Annotation queue for evaluation. I have three suggestions to improve the experience:
Guide users to create prompts with targetable content. I use chat format prompts. My original version mixed static and dynamic content in the user message. Though the dynamic content was delimited with markdown, it was not possible to target it when adding traces to my dataset. I updated my prompt to make the user message content targetable. In this particular case, I had only one variable, so I made the entire user message the dynamic content. This is small thing, but it might be useful to note it in docs/walkthroughs. If it's there already, I missed it.
Guide users to Observations tab for Dataset Additions. This is in the docs, but it might be useful to add something in the UI in or near the Actions menu when a user selects traces. My mind (and maybe Langfuse content?) had me thinking that I would add traces to a dataset, so I was confused when I selected traces and the option was not available. I switched to Observations and found the option I was after. It's a small thing, but it still had me briefly stuck.
Support targeting in Human Annotation Queues. The prompt I'm testing selects "clips" from a portion of a transcript. The ideal way to evaluate this is to see the transcript portion with an indication of which parts were selected/excluded. So far, I've done this by enriching the traces of a dataset run and then reviewing this metadata when processing queue items. It's OK, but not great. I have considered creating a UI specific for this task but would prefer to stay in Langfuse. To improve this experience, consider:
Additional information
No response
All reactions