Skip to content

Releases: s-0-a-r/typesafe-eval

typesafe-eval: v1.3.0

Choose a tag to compare

@github-actions github-actions released this 08 Oct 22:36
bc11fad

1.3.0 (2026-10-08)

Features

  • safety: add Japanese safety fixtures, phone features, and role email patterns (01769ac)
  • safety: add Japanese safety fixtures, phone features, and role email patterns (37e1ac4)

Bug Fixes

  • presets: align tech-spec has_test_plan prompt with labeling criteria to achieve 100% detection (52be2a3)
  • presets: align tech-spec has_test_plan prompt with labeling criteria to achieve 100% detection (dc9d6e3)

Documentation

  • agents: formalize pre-merge /boost review and release DoD checklist (f712a46)
  • agents: formalize pre-merge /boost review and release DoD checklist in AGENTS.md (e7b0c07)
  • restructure README to elevate empirical reliability scope and practical scenarios (19fa5af)
  • restructure README to prioritize empirical scope and practical gating scenarios (385c40d)
  • science: publish empirical accuracy benchmarks for OpenAI Decisions API (a38425f)
  • science: publish empirical accuracy benchmarks for OpenAI Decisions API (1f32614)

typesafe-eval: v1.2.1

Choose a tag to compare

@github-actions github-actions released this 08 Oct 06:06
8ae14de

1.2.1 (2026-10-08)

Bug Fixes

  • science: restore documentation integrity and wire validate multi-provider support (7f7b6e1)
  • science: restore documentation integrity and wire validate multi-provider support (9ba1ced)

typesafe-eval: v1.2.0

Choose a tag to compare

@github-actions github-actions released this 08 Oct 05:39
cbd52ed

1.2.0 (2026-10-08)

Features

  • multimodal: support markdown embedded diagrams via Decisions API (#146) (b8e53cf)
  • multimodal: support markdown embedded diagrams via Decisions API (#146) (0c1068c)

v1.1.0

Choose a tag to compare

@github-actions github-actions released this 07 Oct 23:38
e21e614

1.1.0 (2026-10-07)

Features

  • ac-verification: configure autonomous real CLI acceptance verification system (#94) (f7dc834)
  • agents: Claude Code plugin and AGENTS.md guidance (#35) (#86) (a67942f)
  • baseline: add --baseline diff mode to detect score drops (9502b57)
  • baseline: add --baseline diff mode to detect score drops (closes #39) (1185323)
  • cache: content-addressable evaluation result caching (--cache / --no-cache) (#116) (5687a75), closes #112
  • chunking: chunk long documents for presence questions (#46) (58b7c2c)
  • chunking: chunk long documents for presence questions (Noul) (d3bd671)
  • ci: Phase 4 release automation via release-please, PyPI Trusted Publishing, and Python 3.14 CI (#109) (6a86088)
  • cli: add file exclusion patterns (--exclude), config exclude list, and default ignores (#118) (4df7172)
  • cli: distinct exit codes and continue on file error (0f8f5fe)
  • cli: distinct exit codes and continue on file error (closes #36) (e4c346f)
  • config: support provider and model selection in project config (#141) (3f031cf)
  • docs: setup GitHub Pages deployment workflow and update v1.0.0 doc examples (#130) (b1c78c7)
  • dry-run: show MOCK in table/json/markdown, verdict N/A, exit 0 for every preset (#45) (40be9d2)
  • dry-run: show MOCK in table/json/markdown, verdict N/A, exit 0 for every preset (#45) (541854f)
  • evaluator: extensible preflight binding and transparent override reporting (d5e35ec)
  • evaluator: extensible preflight binding and transparent override reporting (closes #7, closes #8, closes #9, closes #11) (f6e357f)
  • evaluator: support pii preflight override and preserve composite weight for overridden questions (a763c3d)
  • evaluator: support pii preflight override and preserve composite weight for overridden questions (closes #24, closes #25) (53a89f7)
  • initial release of typesafe-eval v0.1.0 (6979f7b)
  • integration: agent-agnostic JSON contract, pre-commit hook and GitHub Action (#34) (#85) (96dcf42)
  • multi-provider decision engine with OpenAI Decisions API support (#139) (#140) (11ca756)
  • offline: offline and rules-only evaluation mode (--offline / --rules-only) (#114) (536153b)
  • openai: align OpenAIDecisionsProvider with verified Decisions API schema (#142) (fe16fe1)
  • position: line and column mapping for PII & secret violations with GitHub Actions annotations (-f github) (#113) (d15c2df)
  • presets: add design_doc and pr_description checklist presets (#41) (1e4e6e1)
  • presets: built-in checklists design_doc and pr_description (#41) (c38d68d)
  • presets: refine confidentiality_risk criteria to 4 levels (Round 1) (5e11efa)
  • presets: remove the two tech-spec Scores that did not track removed content (#43) (ef4803d)
  • presets: round 1 wording for the checklists (d0efa2a)
  • presets: round 1 wording for the design_doc and pr_description checklists (#41) (08a4a26)
  • presets: support custom role_emails glob patterns in preset sanitizer config (255ccd9)
  • presets: support custom role_emails glob patterns in preset sanitizer config (b609a3c)
  • presets: trim tech-spec to has_test_plan and readiness (0a56f28)
  • quality: redefine quality preset as regression detection (#42) (f87de0c)
  • quality: redefine quality preset as regression detection with warnings (#42) (8853c85)
  • reporter: mark scores within 0.1 of threshold as near_threshold (#44) (dd8d28c)
  • reporter: mark values within 0.1 of threshold as near_threshold (#44) (e132645)
  • sanitizer: contextual email PII via state features and dynamic batching (612f1eb)
  • sanitizer: contextual email PII via state features and dynamic batching (closes #30) (e2432e1)
  • sanitizer: per-candidate Noul for phone, ip, url, secrets and decouple detection from masking (c883916)
  • sanitizer: per-candidate Noul for phone, ip, url, secrets and decouple detection from masking (closes #40) (22e585b)
  • skills: add Antigravity code-review skill wi...
Read more

v1.0.0: General Availability (GA)

Choose a tag to compare

@s-0-a-r s-0-a-r released this 06 Oct 05:18
f1f4745

What's Changed in v1.0.0 (GA)

Highlights

  • Official Documentation Portal (Material for MkDocs):
    • Complete, responsive documentation site hosted at GitHub Pages with dark/light themes, search, and syntax highlighting.
    • Comprehensive guides covering Installation, Quickstart, CLI Reference, Python API & Strict Types, CI Integrations (GitHub Actions, pre-commit, Claude Code), Presets, and Scientific Noise Calibration.
    • Strict documentation build validation (mkdocs build --strict) integrated into CI.
  • Empirical Noise Calibration Matrix & Fixture Report:
    • Live empirical evaluation of run-to-run variance on TypeSafe System One (jev-1.13.0) across all 5 built-in presets (60 evaluations, 396 pairwise comparisons).
    • Confirmed 99th percentile $|\Delta| = 0.030$ and max $|\Delta| = 0.040$, validating the statistical robustness of the $\pm 0.10$ threshold safety margin.
    • Automated reusable measurement CLI script (scripts/measure_noise.py) and artifact export (tests/fixtures/noise_report.json).
  • Production Hardening & Verification:
    • 100% pass across all 18 Real CLI Acceptance Criteria (OS subprocess execution matrix).
    • 533 unit, integration, and property-based tests passing with strict typing (PEP 561 py.typed, mypy strict mode).
    • Cleaned up README.md with unified feature catalog and removed temporary release sections.

Real CLI Acceptance Matrix

  • 18/18 Acceptance Criteria verified as OS subprocesses via scripts/verify_ac.py (100% PASS).
  • 533/533 unit, integration, property, and acceptance tests passing across Python 3.10–3.14.

Full Changelog: v0.9.0...v1.0.0

v0.9.0

Choose a tag to compare

@s-0-a-r s-0-a-r released this 06 Oct 00:42
7401735

What's Changed in v0.9.0

Highlights

  • File Exclusion Patterns & Default Ignores (--exclude / -e):
    • Repeatable --exclude <pattern> / -e <pattern> CLI flags supporting wildcard/glob patterns.
    • Project configuration support via exclude: [...] in .typesafe-eval.yaml and pyproject.toml ([tool.typesafe-eval]).
    • Automatic filtering of standard dependency, build, and cache directories (.git, node_modules, .venv, build, dist, etc.) during glob resolution and git diff filtering.
    • Programmatic Python API support via evaluate_documents(paths, exclude=[...]).
  • Offline Rules-Only Mode (--offline / --rules-only):
    • Pure local evaluation relying on deterministic rules (known credential patterns, AWS/Slack/GitHub tokens, free-mail PII) with zero network calls and no required TYPESAFE_API_KEY.
    • Ambiguous candidate fallback behavior producing soft informational notices instead of runtime errors.
    • Programmatic Python API support via evaluate(content, offline=True) and evaluate_documents(paths, offline=True).
  • Content-Addressable Result Caching (--cache / --no-cache):
    • Deterministic SHA-256 caching of evaluation results based on content, preset rules, model parameters, and options.
    • Configurable cache location via --cache-dir or TYPESAFE_CACHE_DIR (default: .typesafe-eval-cache).
    • Cache management CLI command: typesafe-eval cache clear.
    • Programmatic Python API support via evaluate_documents(paths, cache=True, cache_dir=...).
  • Line & Column Mapping with GitHub Actions Annotations (-f github):
    • Exact line and 1-indexed column position tracking for PII, secrets, and structural headings (Position, ViolationItem.line, ViolationItem.col).
    • Native GitHub Actions workflow commands emitted to stdout (::error file=...,line=...,col=...::... and ::warning file=...,line=...,col=...::...) for PR diff inline annotations.
  • Expanded 18-Item Real CLI Acceptance Criteria (AC) Matrix:
    • Full end-to-end OS subprocess execution testing all 18 criteria (scripts/verify_ac.py and pytest -v -m acceptance).
    • Extended coverage for offline mode (AC 15), result cache (AC 16), GitHub Actions annotations (AC 17), and file exclusions (AC 18).

Real CLI Acceptance Matrix

  • 18/18 Acceptance Criteria verified as OS subprocesses via scripts/verify_ac.py (100% PASS).
  • 533/533 unit, integration, property, and acceptance tests passing across Python 3.10–3.14.

Full Changelog: v0.8.0...v0.9.0

v0.8.0

Choose a tag to compare

@s-0-a-r s-0-a-r released this 05 Oct 04:50
5c45920

What's Changed in v0.8.0

Highlights

  • Public Python API: Direct programmatic access via import typesafe_eval (evaluate(), evaluate_document(), evaluate_documents()) with comprehensive typed exceptions (TypeSafeEvalError, ConfigurationError, AuthenticationError, RuntimeEvalError, ContentViolationError).
  • Strict Typing & PEP 561: Packaged src/typesafe_eval/py.typed marker file and zero-error mypy --strict compliance across all 12 source files.
  • Hypothesis Property-Based Testing: Text chunking coverage invariants, secret masking zero-leakage guarantee, and score monotonicity verified in tests/test_properties.py.
  • Configuration Auto-Discovery: Automatic upward directory traversal to locate .typesafe-eval.yaml, .typesafe-eval.yml, or pyproject.toml ([tool.typesafe-eval]), supporting preset extension (extends: quality) and custom rules without mandatory CLI flags.
  • Parallel Document Concurrency: Accelerated multi-document evaluations via thread pooling (-j / --jobs / --concurrency) preserving strict document ordering in outputs.
  • Git Diff Evaluation: Zero-noise CI/pre-commit gating via --staged and --since <ref>.
  • Advisory Thresholds: Soft gate observations via advisory: true questions that inform without tripping Exit Code 1.
  • JSON Schema Export: New CLI command typesafe-eval schema for IDE auto-completion.
  • Adversarial Security Test Suite: 103 adversarial edge-case tests in tests/test_adversarial_sanitizer.py for PII and credential detection.

Real CLI Acceptance Matrix

  • 14/14 Acceptance Criteria verified as OS subprocesses via scripts/verify_ac.py (100% PASS).
  • 494/494 unit, integration, and property tests passing across Python 3.10–3.13.

Full Changelog: v0.7.0...v0.8.0

What's Changed

  • Release v0.8.0: Multi-File Concurrency, Public Python API, Config Discovery & Strict Typing by @s-0-a-r in #107

Full Changelog: v0.7.0...v0.8.0

v0.7.0

Choose a tag to compare

@github-actions github-actions released this 04 Oct 10:47
14cb65e

What's Changed

  • Release v0.7.0: Ruff Adoption, Branch Policy Enforcement, and Lean Package Distribution by @s-0-a-r in #98

Full Changelog: v0.6.0...v0.7.0

v0.6.0

Choose a tag to compare

@github-actions github-actions released this 03 Oct 02:37
f7dc834

What's Changed

  • Release v0.6.0: International Phone PII, Candidate Batching, Baseline Portability, and CLI Init by @s-0-a-r in #92
  • refactor(baseline): prioritize exact normalized match before suffix fallback in cross-environment diff by @s-0-a-r in #93
  • feat: configure autonomous real CLI acceptance verification system by @s-0-a-r in #94

Full Changelog: v0.5.1...v0.6.0

v0.5.1

Choose a tag to compare

@github-actions github-actions released this 02 Oct 02:59
eda580d

What's Changed in v0.5.1

🔒 Security & Sanitizer Hardening

  • JSON/YAML Quoted Key Secret Detection (sanitizer.py): Support secret keys enclosed in single or double quotes (e.g. {"password": "..."}, 'api_key': '...'), preventing secret exposure in JSON payloads and configs.
  • Slack API Token Support: Added deterministic token regex (xox[baprs]-[0-9a-zA-Z-]{10,}) to KNOWN_CREDENTIAL_SPECS.
  • Placeholder Normalization: Registered "changeme" in PLACEHOLDER_SUBSTRINGS.
  • URL Trailing Punctuation Stripping: Trailing punctuation (., ,, ], ), >) is excluded from URL spans and feature extraction, preserving document prose and internal TLD checks.

⚡ Performance & Core Evaluator Optimization

  • Free-dial Prefix Skip: Support and toll-free numbers (0120, 1-800, etc.) skip candidate LLM queries, avoiding token waste and redundant round-trips.
  • Preflight PII Fallback: Added fallback for has_pii in _find_preflight_question.
  • Candidate Spec Ordering: Reordered candidate specs prior to chunking so long documents (> 25k chars) without document-level Nouls properly evaluate candidate questions.

🖥️ CLI & Terminal Reporting

  • Rich Score Badge Color Inversion (reporter.py): Inverted color badges for risk metrics with max_threshold (has_secrets, has_pii) so low risk (0–25%) is green and high risk (>50%) is red.
  • Missing File Handling (cli.py): Explicit non-existent file paths output an error to stderr and exit with code 2.

🤖 CI & Agent Integrations

  • Pre-commit Graceful Skip (hook.py, .pre-commit-hooks.yaml): Added typesafe-eval-hook wrapper to cleanly map Exit 3 (missing API key in pre-commit.ci) to Exit 0.
  • Claude Safety Hook (claude_safety_hook.py): Prevented false blocks (Exit 2) on non-JSON tool output; merged git diff and git status --porcelain for untracked markdown files.
  • Documentation: Corrected repository URL references to s-0-a-r/typesafe-eval.

Full Changelog: v0.5.0...v0.5.1