-
Notifications
You must be signed in to change notification settings - Fork 1
Quick Start
This is the fastest path from a fresh installation to a repeatable evaluation loop, where a rough prompt is tested, tempered, and measured until it holds its shape.
flowchart LR
A[Configure Providers] --> B[Create Prompt]
B --> C[Run Across Models]
C --> D[Create Test Suite]
D --> E[Track Evolution]
Start with one local model and one hosted model if you want a quick read on the balance between quality, cost, and speed.
| Provider | Suggested starter model | Requirement |
|---|---|---|
| Ollama | llama3 |
http://localhost:11434/api/chat available |
| OpenAI | gpt-4o |
OPENAI_API_KEY |
| Anthropic | claude-3-5-sonnet-20241022 |
ANTHROPIC_API_KEY |
See Providers for complete configuration details.
Templates use {{variable}} placeholders so a single prompt can be reused across many inputs without losing its form:
Summarize the following article in 3-5 bullet points:
{{article}}
- Select one or more providers.
- Supply variable values for the template.
- Execute the run and compare outputs side by side.
Each result captures the marks left by the run:
- Response text
- Latency in milliseconds
- Input and output token counts
- Estimated cost in USD
Turn an ad hoc prompt into something you can assay repeatedly:
{
"name": "Phoenix LiveView Summary",
"variable_values": {
"article": "Phoenix LiveView is a library..."
},
"assertions": [
{"type": "contains", "value": "real-time"},
{"type": "regex", "value": "^- .*\\n- .*\\n- .*"}
]
}Supported assertion types:
-
containsfor substring checks -
regexfor pattern validation -
exact_matchfor exact output comparisons -
json_fieldfor structured response validation
Every prompt update creates a new version, which makes it possible to compare:
- Success rate across versions
- Latency trends over time
- Cost changes by provider and prompt version
Use the Evolution view to see whether each change refined the prompt or introduced fresh impurities.
- Prompts for template syntax and versioning
- Providers for provider-specific options
- Architecture for the execution model behind the UI
- Prompts
- Providers
- Runs and Execution
- Evaluation Suites
- Regex Assertions
- Metric Context
- Evaluator Execution Details
- Rubric Judges
- Judge Catalog
- Repeated Sampling
- Quality Policies
- ExUnit Evaluations
- File-Based Suites
- Evaluation Reporters
- Datasets
- Red-Team Datasets
- Generated Red-Team Cases
- Analytics and Prompt Evolution
- Exports and CI
- Documents and Storage
- Embedding and Access