-
Notifications
You must be signed in to change notification settings - Fork 1
Roadmap
Development timeline aligned with the Evaluation Roadmap v1.0
Last Updated: May 15, 2026 — Phase 90 "Adoption First" — monetization on STANDBY
Strategic spine. Per Alton's May 13 ruling (Decision #4), Beta v0.6.0 is now redefined as M0 + M1 from the Evaluation Roadmap — a credibility milestone, NOT a commercial milestone. Pricing is on standby until external evidence justifies it.
Tagged roadmap-v1.0 · Year-one window: May 2026 → April 2027
Read the full roadmap → · Canonical markdown →
| # | Date | Milestone |
|---|---|---|
| M0 | May 2026 | Roadmap published + git tag + CI workflow blocking untagged edits |
| M1 | Jun 2026 | Governance fixes (LLM persona rename, SECURITY.md, MAINTAINERS.md, GOVERNANCE.md, co-maintainer post) |
| M2 | Jul 2026 | Seed labeled eval set v0 — 100 → 200 items on Hugging Face Datasets + Zenodo DOI |
| M3 | Sep 2026 | First Cohen's κ report — three-way (vs human, self-test-retest, cross-family judges) with bootstrap 95% CIs |
| M4 | Oct 2026 | Co-maintainer onboarded or publicly conceded (with retrospective) |
| M5 | Nov 2026 | CS execution sandbox v0 — gVisor / Firecracker, no egress, threat model published |
| M6 | Jan 2027 | Z calibration & abstention — ECE, Brier, reliability diagram against M2 labeled set |
| M7 | Feb 2027 | External benchmark (readiness-gated) — HaluEval / MLCommons AILuminate / SWE-Bench Lite or PaperBench |
| M8 | Mar 2027 | NIST AI RMF self-attestation with external critique published |
| M9 | Apr 2027 | Year-1 retrospective — labeled set at 800 items, decision: continue / pivot / sunset |
Pre-registered thresholds: Cohen's κ ≥ 0.60 (usable claims), κ ≥ 0.80 (production-eligible) · ECE ≤ 0.05 · F1 lift ≥ 0.10 vs single-call baseline · ESR ≥ 0.95 sandbox execution.
Eight pre-registered kill-conditions are stated up front. Failure numbers will ship in the same font size as success numbers.
Released May 15, 2026 — Release v0.5.34 · PR #218
-
/research/evaluation-roadmap— pre-registered Evaluation Roadmap v1.0 live -
Git tag
roadmap-v1.0applied on merge commit — binding mechanism -
Bidirectional cross-link with
/research/paradox -
Canonical markdown at
docs/research/evaluation-roadmap/roadmap-v1.0.md(full Section B technical RFC appendix) - server.json v3.11.0 → MCP Registry republished
-
v0.5.33 (May 13) — Changelog Hygiene: disclosure-policy split between public
/changelog(sanitized) and internalCHANGELOG.md(full forensics) - v0.5.32 (May 13) — Secret Scanner Block (7th IP) + SonarCloud P1 (constants extracted, complexity refactored)
- v0.5.31 (May 13) — SonarCloud P0 cleanup (XV's May 12 audit findings resolved)
- v0.5.28 (May 9) — Tools Free — Option B pivot: paywall removed, all 13 tools free for everyone forever
- v0.5.22 / v0.5.26 / v0.5.30 / v0.5.32 — IP blocklist hardening (7 rogue IPs blocked)
Beta = credibility milestone, NOT commercial milestone.
- M0 (May): Roadmap published ✅
- M1 (June): Governance fixes — LLM persona rename, SECURITY.md (90-day disclosure window), MAINTAINERS.md with open second slot, GOVERNANCE.md (amendment process), co-maintainer search post live
Pricing path is on STANDBY under Phase 90 until external evidence justifies it. The "Early Evaluator" badge replaces outdated tier badges — registration means "help us build the M2 labeled dataset."
Released March 12, 2026
- Creator-centric bias fix: Removed VerifiMind self-promotion from X Agent output
- Added
founder_summaryplain-language layer - Added
research_prompts(Perplexity/Grok bridge)
Released March 10, 2026
- Token Ceiling Monitor: Proactive token usage tracking
- 404 retention fix for improved user experience
- Smithery server-card metadata update
Released March 9, 2026
- Genesis v4.2 "Sentinel-Verified": Forced citations in all agent outputs
- MACP v2.2 "Identity": Multi-Agent Communication Protocol identity model
- L Blind Test: 11/11 perfect score
Released March 7, 2026
- Z-Protocol v1.1: Enhanced ethical review framework
- CS v1.1 "Sentinel": 21 security frameworks, 6-stage pipeline
- OWASP Agentic AI alignment
Released March 1, 2026 — PR #60 | Discussion #62 | Z-Protocol Approved (9.2/10)
-
SessionContext tracing: 8-char
_session_idcorrelation token per Trinity run -
Error handling v2:
build_error_response()for structured, consistent errors -
Health endpoint v2:
health_version: 2with session tracking and BYOK status - Smithery removal: Fully self-hosted on GCP Cloud Run (zero external dependencies)
- BYOK hardening: Retry logic, graceful degradation, provider health checks
- 205 tests, 55.1% coverage: Up from 175 tests at 53.6%
- docs/BYOK_GUIDE.md and docs/SECURITY_SPEC.md shipped
- MIGRATION.md: Smithery → direct connection guide
Released February 28, 2026 — PR #55 | Discussion #58
- Per-tool-call BYOK with ephemeral provider factory
- Auto-detect key format (gsk_ → Groq, sk-ant- → Anthropic, sk- → OpenAI)
- 7+ provider support (Gemini, OpenAI, Anthropic, Groq, Mistral, Ollama, xAI)
- Triple-validated (Manus AI 6/6, Claude Code 6/6, CI 175 tests)
| Version | Name | Target | Key Features |
|---|---|---|---|
| v0.5.x | Stability | Q2 2026 | Bug fixes, performance optimization, monitoring |
| v0.6.0 | Agent Skills | Q2 2026 | Pluggable agent capabilities, shared core engine extraction |
| v0.7.0 | MCP App | Q3 2026 | Full MCP application with persistent sessions |
| v0.8.0+ | Quad Validation | Q3+ 2026 | 4th agent (G-Agent/Grok), separate CLI interface (discussion) |
The FLYWHEEL TEAM has defined 5 conditions that must be met before creating a separate quad-cli repository:
- v0.5.0 Foundation complete and stable ✅
- Proven BYOK adoption (>20 active BYOK users)
- Community demand for CLI interface validated
- Core engine extracted as shared module (v0.6.0)
- Security specification inherited from v0.5.0
Priority chain: v0.5.0 Foundation ✅ → v0.5.x Stability (in progress) → v0.6.0 Agent Skills → v0.7.0 MCP App → v0.8.0+ Quad CLI
| Version | Date | Key Changes |
|---|---|---|
| v0.5.5 | Mar 13, 2026 | Trinity Baseline: TrinitySynthesis schema fix, 3 regression tests, 208 tests |
| v0.5.4 | Mar 12, 2026 | X Agent v4.3: creator-centric bias fix, founder_summary, research_prompts |
| v0.5.3 | Mar 10, 2026 | Token Ceiling Monitor, 404 retention fix, Smithery server-card |
| v0.5.2 | Mar 9, 2026 | Genesis v4.2 "Sentinel-Verified": forced citations, MACP v2.2 "Identity", L Blind Test 11/11 |
| v0.5.1 | Mar 7, 2026 | Z-Protocol v1.1 + CS v1.1 "Sentinel": 21 frameworks, 6-stage, OWASP Agentic AI |
| v0.5.0 | Mar 1, 2026 | Foundation: SessionContext, error handling v2, health v2, Smithery removal, 205 tests |
| v0.4.5 | Feb 28, 2026 | BYOK Live: per-tool-call provider override, auto-detect key format, triple-validated |
| v0.4.4 | Feb 27, 2026 | Multi-Model Trinity (_overall_quality: "full"), X=Gemini, Z/CS=Groq |
| v0.4.3 | Feb 27, 2026 | C-S-P pipeline, system notice, robust JSON extraction |
| v0.4.2 | Feb 26, 2026 | Mock mode resolved, transparent disclosure, CodeQL 13→0 |
| Genesis v3.1 | Feb 2026 | CS Agent 4-Stage Verification Protocol, zero code changes |
| v0.4.1 | Feb 14, 2026 | Markdown-first output, Smithery URL removal, PDF deprecated |
| v0.4.0 | Feb 2026 | Streamable-HTTP transport, GCP Cloud Run migration |
| v0.3.5 | Feb 2026 | Input sanitization, prompt injection detection |
| v0.3.0 | Jan 2026 | X-Z-CS RefleXion Trinity |
| v0.2.0 | Jan 2026 | Template library (19 templates) |
| v0.1.0 | Dec 2025 | Initial MCP server with basic validation |
See the full ROADMAP.md in the repository for detailed planning.
VerifiMind PEAS documentation · Current status · Live health · Public statements · MIT License
Runtime versions, models, routing, tool availability, policies, metrics, and deployment facts are owned by their linked live or release-bound sources.
Start Here
Operator Playbook
Textbook
Trust & Transparency
Evidence & Research
Project Links