You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
First public open-source release of AI Evaluation Platform.
Self-hosted evaluation workbench for RAG, AI Agents, multi-turn conversations, endpoint evaluation, LLM-as-a-Judge, reports, manual review, and human blind testing.
English and Simplified Chinese README entry points with product screenshots.
Included
FastAPI backend with SQLAlchemy models and native evaluation executors.
React 18, TypeScript, Ant Design 5, and Vite frontend.
Dataset management for single-turn rows, multi-turn conversations, RAG contexts, tool calls, reference answers, and custom fields.
Scenario-oriented metrics for RAG, Agent, and multi-turn conversation evaluation.
OpenAI-compatible Judge LLM configuration, with DashScope Qwen Plus as the default example.
Offline evaluation and live endpoint evaluation for saved Chat, RAG, and Agent endpoints.