Skip to content

features evaluator

Claude edited this page Sep 23, 2026 · 1 revision

Evaluator

Active contributors: ferdiiskandar

Purpose

The Evaluator scores a prompt with an LLM-as-a-judge. It asks the provider for a structured rubric response across four dimensions, normalises the scores, and returns a weighted overall score with feedback and improvement suggestions. The desktop console exposes it as evaluate:run and the /evaluate command.

Key abstractions

Abstraction Role Source
evaluatePrompt Entry point: prompt, call provider, parse, score lib/evaluator/engine.ts
DIMENSIONS Four dimensions with weights and rubrics lib/evaluator/dimensions.ts
normalizeScores Clamps to 0–10, fills missing with 5 lib/evaluator/scoring.ts
calculateOverallScore Weighted mean, rounded to one decimal lib/evaluator/scoring.ts
buildEvaluateSystemPrompt JSON rubric contract for the judge lib/llm/prompt-builder.ts

How it works

evaluatePrompt builds the evaluation system and user prompts and calls the provider with maxTokens 2048 and temperature 0.3. The system prompt fixes the output as a JSON object with structure, clarity, completeness, and specificity (each a score 0–10 and a feedback string) plus a suggestions array.

The four dimensions are defined in DIMENSIONS. Each carries a key, label, description, weight, and a 1–10 rubric band. Weights are read from environment variables (EVAL_WEIGHT_STRUCTURE, EVAL_WEIGHT_CLARITY, EVAL_WEIGHT_COMPLETENESS, EVAL_WEIGHT_SPECIFICITY), each defaulting to 0.25. The four values are normalised so they always sum to 1; if the total is not positive, the evaluator falls back to equal quarters.

After the provider responds, the engine strips a markdown fence if present and parses the JSON. normalizeScores maps the parsed object onto the four dimensions in fixed order, clamps each score into 0–10, and substitutes a score of 5 with "No feedback available" for any dimension the model omitted. calculateOverallScore computes the weighted sum divided by the total weight and rounds to one decimal place. getScoreLabel and getScoreColor map the overall score to a band and a colour for display.

If the JSON does not parse, the engine does not fall back to a heuristic. It returns status: 'FAILED' with an empty dimension list and a failure object (code: EVALUATION_PARSE_FAILED, a message telling the user to re-run with the same provider), plus metadata. There is no second attempt.

Integration points

  • Providers and per-surface overrides come from the registry; see LLM providers.
  • The desktop router applies tier quota and model access before calling the engine; see Desktop console.

Entry points for modification

Change weights or rubrics in lib/evaluator/dimensions.ts (or set the environment variables). Change the judge contract in buildEvaluateSystemPrompt in lib/llm/prompt-builder.ts.

Key source files

File Purpose
lib/evaluator/engine.ts Orchestration and failure path
lib/evaluator/dimensions.ts Dimensions, weights, rubrics
lib/evaluator/scoring.ts Normalisation and weighted score
lib/llm/prompt-builder.ts Judge system prompt

Clone this wiki locally