Repository navigation
Releases: s-0-a-r/typesafe-eval
Releases · s-0-a-r/typesafe-eval
Release list
typesafe-eval: v1.3.0
1.3.0 (2026-10-08)
Features
- safety: add Japanese safety fixtures, phone features, and role email patterns (01769ac)
- safety: add Japanese safety fixtures, phone features, and role email patterns (37e1ac4)
Bug Fixes
- presets: align tech-spec has_test_plan prompt with labeling criteria to achieve 100% detection (52be2a3)
- presets: align tech-spec has_test_plan prompt with labeling criteria to achieve 100% detection (dc9d6e3)
Documentation
- agents: formalize pre-merge /boost review and release DoD checklist (f712a46)
- agents: formalize pre-merge /boost review and release DoD checklist in AGENTS.md (e7b0c07)
- restructure README to elevate empirical reliability scope and practical scenarios (19fa5af)
- restructure README to prioritize empirical scope and practical gating scenarios (385c40d)
- science: publish empirical accuracy benchmarks for OpenAI Decisions API (a38425f)
- science: publish empirical accuracy benchmarks for OpenAI Decisions API (1f32614)
typesafe-eval: v1.2.1
typesafe-eval: v1.2.0
v1.1.0
1.1.0 (2026-10-07)
Features
- ac-verification: configure autonomous real CLI acceptance verification system (#94) (f7dc834)
- agents: Claude Code plugin and AGENTS.md guidance (#35) (#86) (a67942f)
- baseline: add --baseline diff mode to detect score drops (9502b57)
- baseline: add --baseline diff mode to detect score drops (closes #39) (1185323)
- cache: content-addressable evaluation result caching (--cache / --no-cache) (#116) (5687a75), closes #112
- chunking: chunk long documents for presence questions (#46) (58b7c2c)
- chunking: chunk long documents for presence questions (Noul) (d3bd671)
- ci: Phase 4 release automation via release-please, PyPI Trusted Publishing, and Python 3.14 CI (#109) (6a86088)
- cli: add file exclusion patterns (--exclude), config exclude list, and default ignores (#118) (4df7172)
- cli: distinct exit codes and continue on file error (0f8f5fe)
- cli: distinct exit codes and continue on file error (closes #36) (e4c346f)
- config: support provider and model selection in project config (#141) (3f031cf)
- docs: setup GitHub Pages deployment workflow and update v1.0.0 doc examples (#130) (b1c78c7)
- dry-run: show MOCK in table/json/markdown, verdict N/A, exit 0 for every preset (#45) (40be9d2)
- dry-run: show MOCK in table/json/markdown, verdict N/A, exit 0 for every preset (#45) (541854f)
- evaluator: extensible preflight binding and transparent override reporting (d5e35ec)
- evaluator: extensible preflight binding and transparent override reporting (closes #7, closes #8, closes #9, closes #11) (f6e357f)
- evaluator: support pii preflight override and preserve composite weight for overridden questions (a763c3d)
- evaluator: support pii preflight override and preserve composite weight for overridden questions (closes #24, closes #25) (53a89f7)
- initial release of typesafe-eval v0.1.0 (6979f7b)
- integration: agent-agnostic JSON contract, pre-commit hook and GitHub Action (#34) (#85) (96dcf42)
- multi-provider decision engine with OpenAI Decisions API support (#139) (#140) (11ca756)
- offline: offline and rules-only evaluation mode (--offline / --rules-only) (#114) (536153b)
- openai: align OpenAIDecisionsProvider with verified Decisions API schema (#142) (fe16fe1)
- position: line and column mapping for PII & secret violations with GitHub Actions annotations (-f github) (#113) (d15c2df)
- presets: add design_doc and pr_description checklist presets (#41) (1e4e6e1)
- presets: built-in checklists design_doc and pr_description (#41) (c38d68d)
- presets: refine confidentiality_risk criteria to 4 levels (Round 1) (5e11efa)
- presets: remove the two tech-spec Scores that did not track removed content (#43) (ef4803d)
- presets: round 1 wording for the checklists (d0efa2a)
- presets: round 1 wording for the design_doc and pr_description checklists (#41) (08a4a26)
- presets: support custom role_emails glob patterns in preset sanitizer config (255ccd9)
- presets: support custom role_emails glob patterns in preset sanitizer config (b609a3c)
- presets: trim tech-spec to has_test_plan and readiness (0a56f28)
- quality: redefine quality preset as regression detection (#42) (f87de0c)
- quality: redefine quality preset as regression detection with warnings (#42) (8853c85)
- reporter: mark scores within 0.1 of threshold as near_threshold (#44) (dd8d28c)
- reporter: mark values within 0.1 of threshold as near_threshold (#44) (e132645)
- sanitizer: contextual email PII via state features and dynamic batching (612f1eb)
- sanitizer: contextual email PII via state features and dynamic batching (closes #30) (e2432e1)
- sanitizer: per-candidate Noul for phone, ip, url, secrets and decouple detection from masking (c883916)
- sanitizer: per-candidate Noul for phone, ip, url, secrets and decouple detection from masking (closes #40) (22e585b)
- skills: add Antigravity code-review skill wi...
v1.0.0: General Availability (GA)
What's Changed in v1.0.0 (GA)
Highlights
-
Official Documentation Portal (Material for MkDocs):
- Complete, responsive documentation site hosted at GitHub Pages with dark/light themes, search, and syntax highlighting.
- Comprehensive guides covering Installation, Quickstart, CLI Reference, Python API & Strict Types, CI Integrations (GitHub Actions, pre-commit, Claude Code), Presets, and Scientific Noise Calibration.
- Strict documentation build validation (
mkdocs build --strict) integrated into CI.
-
Empirical Noise Calibration Matrix & Fixture Report:
- Live empirical evaluation of run-to-run variance on TypeSafe System One (
jev-1.13.0) across all 5 built-in presets (60 evaluations, 396 pairwise comparisons). - Confirmed 99th percentile
$|\Delta| = 0.030$ and max$|\Delta| = 0.040$ , validating the statistical robustness of the$\pm 0.10$ threshold safety margin. - Automated reusable measurement CLI script (
scripts/measure_noise.py) and artifact export (tests/fixtures/noise_report.json).
- Live empirical evaluation of run-to-run variance on TypeSafe System One (
-
Production Hardening & Verification:
- 100% pass across all 18 Real CLI Acceptance Criteria (OS subprocess execution matrix).
- 533 unit, integration, and property-based tests passing with strict typing (PEP 561
py.typed, mypy strict mode). - Cleaned up README.md with unified feature catalog and removed temporary release sections.
Real CLI Acceptance Matrix
- 18/18 Acceptance Criteria verified as OS subprocesses via
scripts/verify_ac.py(100% PASS). - 533/533 unit, integration, property, and acceptance tests passing across Python 3.10–3.14.
Full Changelog: v0.9.0...v1.0.0
v0.9.0
What's Changed in v0.9.0
Highlights
- File Exclusion Patterns & Default Ignores (
--exclude/-e):- Repeatable
--exclude <pattern>/-e <pattern>CLI flags supporting wildcard/glob patterns. - Project configuration support via
exclude: [...]in.typesafe-eval.yamlandpyproject.toml([tool.typesafe-eval]). - Automatic filtering of standard dependency, build, and cache directories (
.git,node_modules,.venv,build,dist, etc.) during glob resolution and git diff filtering. - Programmatic Python API support via
evaluate_documents(paths, exclude=[...]).
- Repeatable
- Offline Rules-Only Mode (
--offline/--rules-only):- Pure local evaluation relying on deterministic rules (known credential patterns, AWS/Slack/GitHub tokens, free-mail PII) with zero network calls and no required
TYPESAFE_API_KEY. - Ambiguous candidate fallback behavior producing soft informational notices instead of runtime errors.
- Programmatic Python API support via
evaluate(content, offline=True)andevaluate_documents(paths, offline=True).
- Pure local evaluation relying on deterministic rules (known credential patterns, AWS/Slack/GitHub tokens, free-mail PII) with zero network calls and no required
- Content-Addressable Result Caching (
--cache/--no-cache):- Deterministic SHA-256 caching of evaluation results based on content, preset rules, model parameters, and options.
- Configurable cache location via
--cache-dirorTYPESAFE_CACHE_DIR(default:.typesafe-eval-cache). - Cache management CLI command:
typesafe-eval cache clear. - Programmatic Python API support via
evaluate_documents(paths, cache=True, cache_dir=...).
- Line & Column Mapping with GitHub Actions Annotations (
-f github):- Exact line and 1-indexed column position tracking for PII, secrets, and structural headings (
Position,ViolationItem.line,ViolationItem.col). - Native GitHub Actions workflow commands emitted to
stdout(::error file=...,line=...,col=...::...and::warning file=...,line=...,col=...::...) for PR diff inline annotations.
- Exact line and 1-indexed column position tracking for PII, secrets, and structural headings (
- Expanded 18-Item Real CLI Acceptance Criteria (AC) Matrix:
- Full end-to-end OS subprocess execution testing all 18 criteria (
scripts/verify_ac.pyandpytest -v -m acceptance). - Extended coverage for offline mode (AC 15), result cache (AC 16), GitHub Actions annotations (AC 17), and file exclusions (AC 18).
- Full end-to-end OS subprocess execution testing all 18 criteria (
Real CLI Acceptance Matrix
- 18/18 Acceptance Criteria verified as OS subprocesses via
scripts/verify_ac.py(100% PASS). - 533/533 unit, integration, property, and acceptance tests passing across Python 3.10–3.14.
Full Changelog: v0.8.0...v0.9.0
v0.8.0
What's Changed in v0.8.0
Highlights
- Public Python API: Direct programmatic access via
import typesafe_eval(evaluate(),evaluate_document(),evaluate_documents()) with comprehensive typed exceptions (TypeSafeEvalError,ConfigurationError,AuthenticationError,RuntimeEvalError,ContentViolationError). - Strict Typing & PEP 561: Packaged
src/typesafe_eval/py.typedmarker file and zero-errormypy --strictcompliance across all 12 source files. - Hypothesis Property-Based Testing: Text chunking coverage invariants, secret masking zero-leakage guarantee, and score monotonicity verified in
tests/test_properties.py. - Configuration Auto-Discovery: Automatic upward directory traversal to locate
.typesafe-eval.yaml,.typesafe-eval.yml, orpyproject.toml([tool.typesafe-eval]), supporting preset extension (extends: quality) and custom rules without mandatory CLI flags. - Parallel Document Concurrency: Accelerated multi-document evaluations via thread pooling (
-j/--jobs/--concurrency) preserving strict document ordering in outputs. - Git Diff Evaluation: Zero-noise CI/pre-commit gating via
--stagedand--since <ref>. - Advisory Thresholds: Soft gate observations via
advisory: truequestions that inform without tripping Exit Code 1. - JSON Schema Export: New CLI command
typesafe-eval schemafor IDE auto-completion. - Adversarial Security Test Suite: 103 adversarial edge-case tests in
tests/test_adversarial_sanitizer.pyfor PII and credential detection.
Real CLI Acceptance Matrix
- 14/14 Acceptance Criteria verified as OS subprocesses via
scripts/verify_ac.py(100% PASS). - 494/494 unit, integration, and property tests passing across Python 3.10–3.13.
Full Changelog: v0.7.0...v0.8.0
What's Changed
- Release v0.8.0: Multi-File Concurrency, Public Python API, Config Discovery & Strict Typing by @s-0-a-r in #107
Full Changelog: v0.7.0...v0.8.0
v0.7.0
What's Changed
- Release v0.7.0: Ruff Adoption, Branch Policy Enforcement, and Lean Package Distribution by @s-0-a-r in #98
Full Changelog: v0.6.0...v0.7.0
v0.6.0
What's Changed
- Release v0.6.0: International Phone PII, Candidate Batching, Baseline Portability, and CLI Init by @s-0-a-r in #92
- refactor(baseline): prioritize exact normalized match before suffix fallback in cross-environment diff by @s-0-a-r in #93
- feat: configure autonomous real CLI acceptance verification system by @s-0-a-r in #94
Full Changelog: v0.5.1...v0.6.0
v0.5.1
What's Changed in v0.5.1
🔒 Security & Sanitizer Hardening
- JSON/YAML Quoted Key Secret Detection (
sanitizer.py): Support secret keys enclosed in single or double quotes (e.g.{"password": "..."},'api_key': '...'), preventing secret exposure in JSON payloads and configs. - Slack API Token Support: Added deterministic token regex (
xox[baprs]-[0-9a-zA-Z-]{10,}) toKNOWN_CREDENTIAL_SPECS. - Placeholder Normalization: Registered
"changeme"inPLACEHOLDER_SUBSTRINGS. - URL Trailing Punctuation Stripping: Trailing punctuation (
.,,,],),>) is excluded from URL spans and feature extraction, preserving document prose and internal TLD checks.
⚡ Performance & Core Evaluator Optimization
- Free-dial Prefix Skip: Support and toll-free numbers (
0120,1-800, etc.) skip candidate LLM queries, avoiding token waste and redundant round-trips. - Preflight PII Fallback: Added fallback for
has_piiin_find_preflight_question. - Candidate Spec Ordering: Reordered candidate specs prior to chunking so long documents (> 25k chars) without document-level Nouls properly evaluate candidate questions.
🖥️ CLI & Terminal Reporting
- Rich Score Badge Color Inversion (
reporter.py): Inverted color badges for risk metrics withmax_threshold(has_secrets,has_pii) so low risk (0–25%) is green and high risk (>50%) is red. - Missing File Handling (
cli.py): Explicit non-existent file paths output an error tostderrand exit with code 2.
🤖 CI & Agent Integrations
- Pre-commit Graceful Skip (
hook.py,.pre-commit-hooks.yaml): Addedtypesafe-eval-hookwrapper to cleanly map Exit 3 (missing API key inpre-commit.ci) to Exit 0. - Claude Safety Hook (
claude_safety_hook.py): Prevented false blocks (Exit 2) on non-JSON tool output; mergedgit diffandgit status --porcelainfor untracked markdown files. - Documentation: Corrected repository URL references to
s-0-a-r/typesafe-eval.
Full Changelog: v0.5.0...v0.5.1