2.4.0
🧠 Jev integration
SuperQode 2.4.0 adds a general decision harness for Jev, native tool-permission checks, labelled decision evaluations, and opt-in rubric grading.
Coding models keep generating code. Jev supplies typed decisions that SuperQode validates and applies through explicit policies.
✨ Added
-
🧩 General decision harness
Thesystemonebackend evaluates reviewed Choice, Score, and Noul question packs. Input schemas, confidence thresholds, abstention, and pack hashes make each decision inspectable and reproducible. Examples cover ticket triage, factory-route suggestions, and tool permissions. -
🛡️ Native tool-permission checks
:systemone liveenables Jev checks in Core/BYOK sessions. Hard policy denials stay authoritative. Model ASK uses the approval flow even for otherwise auto-allowed tools. Client failures are visible and fall back to the existing permission policy. -
🔌 Direct decision sessions
:systemone connect <pack>and:connect systemone <pack>connect the TUI directly to Jev without a coding provider. Status distinguishes a decision session from the tool-check sidecar. -
📊 Labelled evaluations
harness evalsupports exact typed output labels and explicit evaluator results. Scorecards include coverage, accuracy among graded cases, abstentions, errors, dataset hashes, and decision evidence. Synthetic routing and tool-permission starter datasets include held-in and held-out splits. -
📝 Rubric grading
SUPERQODE_RUBRIC_GRADER=systemoneuses Jev verdicts in the native rubric revision loop. Ajev_rubrictask evaluator judges another harness's response. Headless JSON exposesrubric_result. -
📡 Transport and observability
Live decisions report returned model, HTTP status, latency, and available token usage. Configurable compatible endpoints, stub clients, sanitized recordings, and replay support development and testing.
🔄 Changed
- Invalid, uncertain, or unavailable rubric judgments are explicitly
ungraded. Utility-grader errors no longer count assatisfied. Headless runs with an unsatisfied or ungraded rubric reportsuccess: falseand exit with code2. - Legacy evaluation tasks keep their non-empty smoke check, now identified as
smoke_only. Use labelled or rubric evaluators for stronger assessment.
🔒 Reliability
- Tool-gate pack 1.1.0 requires low destructive and exfiltration risk before auto-approval. Uncertain risk cannot produce ALLOW.
- Live criteria serialization matches the API contract. Evaluation deadlines include retries, and recording failures do not discard valid decisions.
- Input redaction covers common credential fields, headers, command flags, URL credentials, and private-key blocks. Offline policies remain respected.
📌 Scope
Jev integration is opt-in and needs separate TypeSafe credentials for live use.
- The native gate does not intercept external agent runtimes.
- Route decisions are suggestions; automatic factory routing and automatic LLM fallback are not included.
- Starter datasets are examples, not published accuracy benchmarks.
See the Jev integration guide and 2.4.0 release notes for setup and migration.
Install: curl -fsSL https://superqode.dev/install.sh | sh
Update: superqode update or uv tool upgrade superqode
🌐 https://superqode.dev
📦 https://github.com/SuperagenticAI/superqode/releases/tag/v2.4.0