Repository navigation
features evaluator
Active contributors: ferdiiskandar
The Evaluator scores a prompt with an LLM-as-a-judge. It asks the provider for a structured rubric response across four dimensions, normalises the scores, and returns a weighted overall score with feedback and improvement suggestions. The desktop console exposes it as evaluate:run and the /evaluate command.
| Abstraction | Role | Source |
|---|---|---|
evaluatePrompt |
Entry point: prompt, call provider, parse, score | lib/evaluator/engine.ts |
DIMENSIONS |
Four dimensions with weights and rubrics | lib/evaluator/dimensions.ts |
normalizeScores |
Clamps to 0–10, fills missing with 5 | lib/evaluator/scoring.ts |
calculateOverallScore |
Weighted mean, rounded to one decimal | lib/evaluator/scoring.ts |
buildEvaluateSystemPrompt |
JSON rubric contract for the judge | lib/llm/prompt-builder.ts |
evaluatePrompt builds the evaluation system and user prompts and calls the provider with maxTokens 2048 and temperature 0.3. The system prompt fixes the output as a JSON object with structure, clarity, completeness, and specificity (each a score 0–10 and a feedback string) plus a suggestions array.
The four dimensions are defined in DIMENSIONS. Each carries a key, label, description, weight, and a 1–10 rubric band. Weights are read from environment variables (EVAL_WEIGHT_STRUCTURE, EVAL_WEIGHT_CLARITY, EVAL_WEIGHT_COMPLETENESS, EVAL_WEIGHT_SPECIFICITY), each defaulting to 0.25. The four values are normalised so they always sum to 1; if the total is not positive, the evaluator falls back to equal quarters.
After the provider responds, the engine strips a markdown fence if present and parses the JSON. normalizeScores maps the parsed object onto the four dimensions in fixed order, clamps each score into 0–10, and substitutes a score of 5 with "No feedback available" for any dimension the model omitted. calculateOverallScore computes the weighted sum divided by the total weight and rounds to one decimal place. getScoreLabel and getScoreColor map the overall score to a band and a colour for display.
If the JSON does not parse, the engine does not fall back to a heuristic. It returns status: 'FAILED' with an empty dimension list and a failure object (code: EVALUATION_PARSE_FAILED, a message telling the user to re-run with the same provider), plus metadata. There is no second attempt.
- Providers and per-surface overrides come from the registry; see LLM providers.
- The desktop router applies tier quota and model access before calling the engine; see Desktop console.
Change weights or rubrics in lib/evaluator/dimensions.ts (or set the environment variables). Change the judge contract in buildEvaluateSystemPrompt in lib/llm/prompt-builder.ts.
| File | Purpose |
|---|---|
lib/evaluator/engine.ts |
Orchestration and failure path |
lib/evaluator/dimensions.ts |
Dimensions, weights, rubrics |
lib/evaluator/scoring.ts |
Normalisation and weighted score |
lib/llm/prompt-builder.ts |
Judge system prompt |
Open-source prompt engineering and multi-LLM tooling by Sentra Artificial Intelligence.
MyPrompt develops practical approaches to prompt engineering, multi-LLM optimisation, reusable prompt systems, and AI-native workflows — with an emphasis on structured, interoperable, and real-world AI use.
Built in Indonesia as part of the Sentra Artificial Intelligence ecosystem.
Sentra Artificial Intelligence · Source Repository · Official Website
Dr Ferdi Iskandar — Creator & Maintainer
LinkedIn ·
ORCID ·
Hugging Face ·
Kaggle ·
Medium ·
Substack ·
X ·
Threads
MyPrompt · Sentra Artificial Intelligence · Indonesia