Langfuse support for Audio prompt input / output #11740
Describe your questionIs there Langfuse support for speech-to-text / text-to-speech models? Is there documentation to demonstrate speech-to-text / text-to-speech implementation? Langfuse Cloud or Self-Hosted?Self-Hosted If Self-Hostedv3.148.0 If Langfuse CloudNo response SDK and integration versionsNo response Pre-Submission Checklist
|
Replies: 2 comments 2 replies
|
Yes, Langfuse supports audio input and output for speech-to-text and text-to-speech models(1). Audio Input and Output SupportLangfuse provides multi-modal tracing capabilities that include audio attachments(1). The platform can handle audio data through the Documentation and ExamplesThe multi-modality documentation demonstrates audio implementation with the OpenAI SDK(1). Here's an example showing audio input and output: from langfuse import get_client
content_path = "static/joke_prompt.wav"
base64_string = encode_file(content_path)
response = client.chat.completions.create(
model="gpt-4o-audio-preview",
modalities=["text", "audio"],
audio={"voice": "alloy", "format": "wav"},
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Do what this recording says."},
{
"type": "input_audio",
"input_audio": {"data": base64_string, "format": "wav"},
},
],
},
],
)
langfuse = get_client()
langfuse.flush()Cost Tracking for TTS ModelsFor text-to-speech cost tracking, you need to log TTS runs as To enable cost tracking for TTS:
The usage detail keys in your code must exactly match the usage type keys in your Langfuse model definition for cost calculation to work(2). 📚 Sources: Have another question? Just tag @inkeep. |
|
Based on the available sources, I don't have specific information about how the Langfuse UI supports end-user prompt engineering for speech-to-text or text-to-speech modalities. The sources show that Langfuse supports multi-modal tracing with audio attachments(1), and you can log audio input/output in traces programmatically(1). However, the documentation doesn't describe UI features specifically designed for prompt engineering with audio modalities. For general prompt management, there is a discussion about enhancing the Langfuse UI to support tool calls and agent configurations(2), but this doesn't address audio-specific prompt engineering capabilities. If you need to work with audio prompts, you would currently need to:
For more detailed information about UI support for audio prompt engineering, I recommend checking the latest Langfuse documentation or reaching out to the Langfuse team directly, as this specific capability may not be fully documented in the available sources. ** 📚 Sources:
Have another question? Just tag @inkeep. |
@godclause, no, the Langfuse Playground and Langfuse in-UI Experiments do not support audio yet. We have this on our roadmap!