Skip to content

v0.6.1 — Hardening Release

Choose a tag to compare

@JDeun JDeun released this 25 Apr 05:06
· 192 commits to main since this release

Helm v0.6.1 — Hardening Release

Date: 2026-04-25

Helm 0.6.1 is a comprehensive hardening release that improves every module's type safety, immutability, thread safety, and test coverage. No new features are added — this release focuses entirely on correctness and robustness. 118 new tests bring the total to 298.

Hardening Highlights

Security (command_guard, run_with_profile)

  • SemanticResult NamedTuple: Replaced stringly-typed "approve."/"deny." prefix convention with a structured SemanticResult(action, reason) return type
  • Recursive shell unwrapping: _effective_argv loops up to depth 5 to catch nested bash -c "bash -c '...'" patterns
  • dd read/write distinction: if= device → require_approval, of= device → deny
  • shred/wipefs/blkdiscard: Added to semantic deny rules for device-targeting commands
  • Subprocess timeout: --timeout flag (default 1800s) with proper TimeoutExpired handling
  • Minimal environment: _minimal_env() strips secrets from restricted profile subprocesses
  • --guard-json: Machine-readable guard decision output without executing the command
  • Fail-closed tuples: All fallback guard decisions use immutable tuple() instead of list literals

Immutability Enforcement

  • Frozen dataclasses everywhere: All list[str] fields → tuple[str, ...] across GuardDecision, CommandClassification, ProviderProbe, RuntimeFingerprint, HardwareProfile, RuntimeModelState
  • StrategyConfig: New frozen dataclass replaces the lone mutable dict[str, object] field in DiscoverySnapshot
  • IntelligenceTier.available_tiers(): Returns tuple[str, ...] instead of list[str]

Thread Safety & Reliability

  • ops_db: _INITIALIZED_DBS protected by threading.Lock; schema versioning via _check_schema_version()
  • state_io: Windows sentinel-region lock (bytes 0–1); threading.Event for lock warning; documented "ab" mode seek behavior
  • GPU detection: @functools.lru_cache(maxsize=1) on _detect_gpu() — hardware doesn't change during process lifetime
  • Streaming JSONL: verify_index and read_jsonl use line-by-line reading (no OOM on large files)
  • Response body limit: model_provider_probe caps HTTP response reads at 64KB

Code Quality

  • JSONL consolidation: Duplicate read_jsonl/append_jsonl removed from 6 modules → single canonical source in commands/__init__.py and state_io.py
  • run_script consolidation: Removed from 6 command modules → single commands/__init__.py
  • Lazy initialization: reply_gate and run_with_profile defer workspace lookups
  • sys.executable: Replaces hardcoded python3 for Windows compatibility
  • Deep merge: _deep_merge(base, overlay) for skill contract resolution
  • intelligence_tier: Complete rewrite from stub to snapshot-driven L0-L4 resolution

Tests

  • 118 new tests (298 total, was 180)
  • New test suites: test_intelligence_tier.py (19), test_reply_gate.py (19), test_memory_capture.py (23)
  • Expanded: test_command_guard.py (+35), test_run_with_profile_guard.py (+13)

Review Score Progression

Cycle Average CRITICAL IMPORTANT
1st 6.2/10 5 12
2nd 7.7/10 0 12
3rd 8.5/10 0 0
4th 8.8/10 0 0