How to Evaluate an External API #10341
Unanswered
ayub-kokabi
asked this question in
Support
Replies: 2 comments 1 reply
|
To evaluate your external translation API with Langfuse and track the impact of prompt changes, follow these best practices using the built-in features: 1. Prepare Your DatasetYou already have a dataset named Sample in Langfuse. Ensure:
2. Create or Update Your PromptIn Langfuse Prompt Management, make sure your prompt variables match the dataset item keys. For your use case, you’ll want variables like 3. Run Your API and Log Outputs to LangfuseFor each dataset item:
4. Set Up LLM-as-a-Judge Evaluator
5. Configure and Run Prompt Experiment via UI
6. Analyze Results
Tips & Notes
References for Deep Dives📚 Sources:
Have another question? Just tag @inkeep. |
0 replies
|
If someone takes the time to post their question here and patiently waits for a human response, there's a good chance—about 99%—that they've already tried asking AI chatbots! |
1 reply
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
I have a translation API that works in the background using a prompt and leveraging OpenRouter to translate user text:
The response is a JSON object with this structure:
{ "result": "سلام دنیا" }I have a dataset named Sample in the Langfuse datasets section, which has two columns:
{ "content":"How are you?", "to": "fa" }Now I want to evaluate this service using Langfuse.
That is, every time I change the background prompt of the product, I want it to run an evaluation on the dataset I have already prepared.
The model
google/gemini-2.0-flash-lite-001on OpenRouter will act as a judge, taking the API translation output along with the expected output and the original text, and give a score based on a prompt. This way, I can see whether the performance has improved, worsened, or stayed the same.Prompt:
{ "score": <0-1>, "feedback": "<short explanation>" }The more I read Langfuse documentation, the more confused I get.
Right now, I don’t know how to use the Prompts section, LLM as a judge, Run experiment via user interface, etc., to actually perform this evaluation.
What are the best practices for evaluating this system?
I also have more complex scenarios, but I wanted to start with this one first to get into Langfuse and then move on to those sections.
All reactions