Skip to content

Roadmap

creator35lwb-web edited this page May 15, 2026 · 10 revisions

Roadmap

Development timeline aligned with the Evaluation Roadmap v1.0

Last Updated: May 15, 2026 — Phase 90 "Adoption First" — monetization on STANDBY

Strategic spine. Per Alton's May 13 ruling (Decision #4), Beta v0.6.0 is now redefined as M0 + M1 from the Evaluation Roadmap — a credibility milestone, NOT a commercial milestone. Pricing is on standby until external evidence justifies it.


🆕 The Evaluation Roadmap v1.0 — Pre-Registered Milestones

Tagged roadmap-v1.0 · Year-one window: May 2026 → April 2027

Read the full roadmap → · Canonical markdown →

# Date Milestone
M0 May 2026 Roadmap published + git tag + CI workflow blocking untagged edits
M1 Jun 2026 Governance fixes (LLM persona rename, SECURITY.md, MAINTAINERS.md, GOVERNANCE.md, co-maintainer post)
M2 Jul 2026 Seed labeled eval set v0 — 100 → 200 items on Hugging Face Datasets + Zenodo DOI
M3 Sep 2026 First Cohen's κ report — three-way (vs human, self-test-retest, cross-family judges) with bootstrap 95% CIs
M4 Oct 2026 Co-maintainer onboarded or publicly conceded (with retrospective)
M5 Nov 2026 CS execution sandbox v0 — gVisor / Firecracker, no egress, threat model published
M6 Jan 2027 Z calibration & abstention — ECE, Brier, reliability diagram against M2 labeled set
M7 Feb 2027 External benchmark (readiness-gated) — HaluEval / MLCommons AILuminate / SWE-Bench Lite or PaperBench
M8 Mar 2027 NIST AI RMF self-attestation with external critique published
M9 Apr 2027 Year-1 retrospective — labeled set at 800 items, decision: continue / pivot / sunset

Pre-registered thresholds: Cohen's κ ≥ 0.60 (usable claims), κ ≥ 0.80 (production-eligible) · ECE ≤ 0.05 · F1 lift ≥ 0.10 vs single-call baseline · ESR ≥ 0.95 sandbox execution.

Eight pre-registered kill-conditions are stated up front. Failure numbers will ship in the same font size as success numbers.


Current: v0.5.34 — Evaluation Roadmap v1.0 ✅ LIVE

Released May 15, 2026 — Release v0.5.34 · PR #218

  • /research/evaluation-roadmap — pre-registered Evaluation Roadmap v1.0 live
  • Git tag roadmap-v1.0 applied on merge commit — binding mechanism
  • Bidirectional cross-link with /research/paradox
  • Canonical markdown at docs/research/evaluation-roadmap/roadmap-v1.0.md (full Section B technical RFC appendix)
  • server.json v3.11.0 → MCP Registry republished

Recent v0.5.x releases

  • v0.5.33 (May 13) — Changelog Hygiene: disclosure-policy split between public /changelog (sanitized) and internal CHANGELOG.md (full forensics)
  • v0.5.32 (May 13) — Secret Scanner Block (7th IP) + SonarCloud P1 (constants extracted, complexity refactored)
  • v0.5.31 (May 13) — SonarCloud P0 cleanup (XV's May 12 audit findings resolved)
  • v0.5.28 (May 9) — Tools Free — Option B pivot: paywall removed, all 13 tools free for everyone forever
  • v0.5.22 / v0.5.26 / v0.5.30 / v0.5.32 — IP blocklist hardening (7 rogue IPs blocked)

Next: v0.6.0-Beta — Evaluation Roadmap M0 + M1 (June 2026)

Beta = credibility milestone, NOT commercial milestone.

  • M0 (May): Roadmap published ✅
  • M1 (June): Governance fixes — LLM persona rename, SECURITY.md (90-day disclosure window), MAINTAINERS.md with open second slot, GOVERNANCE.md (amendment process), co-maintainer search post live

Pricing path is on STANDBY under Phase 90 until external evidence justifies it. The "Early Evaluator" badge replaces outdated tier badges — registration means "help us build the M2 labeled dataset."


v0.5.4 — X Agent v4.3 ✅

Released March 12, 2026

  • Creator-centric bias fix: Removed VerifiMind self-promotion from X Agent output
  • Added founder_summary plain-language layer
  • Added research_prompts (Perplexity/Grok bridge)

v0.5.3 — Token Ceiling Monitor ✅

Released March 10, 2026

  • Token Ceiling Monitor: Proactive token usage tracking
  • 404 retention fix for improved user experience
  • Smithery server-card metadata update

v0.5.2 — Genesis v4.2 "Sentinel-Verified" ✅

Released March 9, 2026

  • Genesis v4.2 "Sentinel-Verified": Forced citations in all agent outputs
  • MACP v2.2 "Identity": Multi-Agent Communication Protocol identity model
  • L Blind Test: 11/11 perfect score

v0.5.1 — Z-Protocol v1.1 + CS v1.1 "Sentinel" ✅

Released March 7, 2026

  • Z-Protocol v1.1: Enhanced ethical review framework
  • CS v1.1 "Sentinel": 21 security frameworks, 6-stage pipeline
  • OWASP Agentic AI alignment

v0.5.0 — Foundation ✅

Released March 1, 2026 — PR #60 | Discussion #62 | Z-Protocol Approved (9.2/10)

  • SessionContext tracing: 8-char _session_id correlation token per Trinity run
  • Error handling v2: build_error_response() for structured, consistent errors
  • Health endpoint v2: health_version: 2 with session tracking and BYOK status
  • Smithery removal: Fully self-hosted on GCP Cloud Run (zero external dependencies)
  • BYOK hardening: Retry logic, graceful degradation, provider health checks
  • 205 tests, 55.1% coverage: Up from 175 tests at 53.6%
  • docs/BYOK_GUIDE.md and docs/SECURITY_SPEC.md shipped
  • MIGRATION.md: Smithery → direct connection guide

v0.4.5 — BYOK Live ✅

Released February 28, 2026 — PR #55 | Discussion #58

  • Per-tool-call BYOK with ephemeral provider factory
  • Auto-detect key format (gsk_ → Groq, sk-ant- → Anthropic, sk- → OpenAI)
  • 7+ provider support (Gemini, OpenAI, Anthropic, Groq, Mistral, Ollama, xAI)
  • Triple-validated (Manus AI 6/6, Claude Code 6/6, CI 175 tests)

Planned Releases

Version Name Target Key Features
v0.5.x Stability Q2 2026 Bug fixes, performance optimization, monitoring
v0.6.0 Agent Skills Q2 2026 Pluggable agent capabilities, shared core engine extraction
v0.7.0 MCP App Q3 2026 Full MCP application with persistent sessions
v0.8.0+ Quad Validation Q3+ 2026 4th agent (G-Agent/Grok), separate CLI interface (discussion)

Commercialization Path

The FLYWHEEL TEAM has defined 5 conditions that must be met before creating a separate quad-cli repository:

  1. v0.5.0 Foundation complete and stable ✅
  2. Proven BYOK adoption (>20 active BYOK users)
  3. Community demand for CLI interface validated
  4. Core engine extracted as shared module (v0.6.0)
  5. Security specification inherited from v0.5.0

Priority chain: v0.5.0 Foundation ✅ → v0.5.x Stability (in progress) → v0.6.0 Agent Skills → v0.7.0 MCP App → v0.8.0+ Quad CLI


Complete Version History

Version Date Key Changes
v0.5.5 Mar 13, 2026 Trinity Baseline: TrinitySynthesis schema fix, 3 regression tests, 208 tests
v0.5.4 Mar 12, 2026 X Agent v4.3: creator-centric bias fix, founder_summary, research_prompts
v0.5.3 Mar 10, 2026 Token Ceiling Monitor, 404 retention fix, Smithery server-card
v0.5.2 Mar 9, 2026 Genesis v4.2 "Sentinel-Verified": forced citations, MACP v2.2 "Identity", L Blind Test 11/11
v0.5.1 Mar 7, 2026 Z-Protocol v1.1 + CS v1.1 "Sentinel": 21 frameworks, 6-stage, OWASP Agentic AI
v0.5.0 Mar 1, 2026 Foundation: SessionContext, error handling v2, health v2, Smithery removal, 205 tests
v0.4.5 Feb 28, 2026 BYOK Live: per-tool-call provider override, auto-detect key format, triple-validated
v0.4.4 Feb 27, 2026 Multi-Model Trinity (_overall_quality: "full"), X=Gemini, Z/CS=Groq
v0.4.3 Feb 27, 2026 C-S-P pipeline, system notice, robust JSON extraction
v0.4.2 Feb 26, 2026 Mock mode resolved, transparent disclosure, CodeQL 13→0
Genesis v3.1 Feb 2026 CS Agent 4-Stage Verification Protocol, zero code changes
v0.4.1 Feb 14, 2026 Markdown-first output, Smithery URL removal, PDF deprecated
v0.4.0 Feb 2026 Streamable-HTTP transport, GCP Cloud Run migration
v0.3.5 Feb 2026 Input sanitization, prompt injection detection
v0.3.0 Jan 2026 X-Z-CS RefleXion Trinity
v0.2.0 Jan 2026 Template library (19 templates)
v0.1.0 Dec 2025 Initial MCP server with basic validation

See the full ROADMAP.md in the repository for detailed planning.


← Back to Home

Clone this wiki locally