Replies: 1 comment 2 replies
|
Based on the available documentation, you can partially achieve your workflow using existing Langfuse features, though with some limitations: What you can do now:
Current limitations: The main challenge is that Langfuse's current architecture expects actual model runs and API calls. The documentation shows that prompt experiments and evaluations are designed around real LLM providers(3). There's an open discussion about supporting custom model adapters(4), where a user requested similar functionality for custom applications that don't directly access LLMs(4). The suggested workaround was to "expose your application in an OpenAI schema-compatible format"(4), but this still requires some form of API endpoint. Potential workaround: You could manually create traces in Langfuse using the SDKs and populate them with your custom model names and manually generated outputs, then use the annotation features to evaluate them. However, this would require some custom implementation work on your part. For your specific use case of completely offline evaluation with custom model names and no API calls, you might need to wait for enhanced custom adapter support or implement a workaround using the existing SDK to manually populate traces. 📚 Sources:
Have another question? Just tag @inkeep. |
Uh oh!
There was an error while loading. Please reload this page.
Hi, I want to evaluate different coding agents with different models and create/annotate and evaluate the output manually without any actual model runs or api calls.
I want to give codex/claude code/droid/opencode, etc... any permutation of models for each agent the same prompt and run them > evaluate manually and annotate if they did what I asked + annotate some subjective outcome ( pretty ui or bad ui, etc) > view reporting results after my experiments with different permutations of agent/model + prompt versions.
Basically I want an option to create a custom model name ( not a full working setup, no working api). For example "codex/gpt-5-codex-mini", "codex/gpt-5-codex-high", "cc/opus-4-5" and give them "prompt20-v-1" and "prompt20-v-2" and without running them through promptfoo annotate the result and even skip the output completely or fill it manually.
Is there a way to do it now without creating any new features?
Thanks
All reactions