Skip to content

Release 0.6.9

Choose a tag to compare

@StevenBtw StevenBtw released this 01 Mar 10:35
· 68 commits to main since this release
832962d

v0.6.9 - Benchmarks, Relationships & Analysis (March 1, 2026)

I have been running a lot of benchmarks on flask_invoice_generator, full-stack-fastapi-template and taiga-back/taiga-front. Besides a lot of new config versions, I also added a few improvements to further reduce tokens per run and added some things to make my life easier while benchmarking and trying to get the % up without any overly canonical or repository specific prompts. Currently not breaking the 60% barrier, for relationships (the hardest one).

CLI

  • Single Step Execution: New --only-step option for run command to run a single extraction/derivation step (disables all others)
  • Benchmark Step Isolation: New --only-extraction-step and --only-derivation-step options for benchmark runs
  • Enrichment Cache Control: New --nocache-enrichment-configs option for selective cache bypass during benchmarks
  • Sequence Reordering: New config sequence command to reorder derivation step execution (e.g., bottom-up: Technology → Application → Business)
  • Read-Only Config Access: New config query command for safe config access during benchmark runs (non-blocking)
  • Batch Size in CLI: Added --batch-size option to config update for extraction batching

Extraction

  • Edge Extraction Module: New edges.py for Tree-sitter based relationship extraction. Extracts IMPORTS, USES, CALLS, DECORATED_BY, and REFERENCES edges in a single efficient parse per file. Language-specific filter constants for Python, JavaScript, Java, and C#. Fixed type node ID format mismatch that caused REFERENCES edge creation failures
  • Directory Classification Step: New extraction step after directories to create technology and business concept nodes (with batched LLM calls), guiding subsequent LLM extraction
  • Structural Technology Extraction: New extraction method for Technology nodes from infrastructure files (docker-compose.yml, Dockerfile, .env) without LLM
  • Token Efficiency: Compact JSON serialization (~15% savings), system/user prompt separation, and multi-file batching (--batch-size N). Estimated 40-60% total reduction
  • Error Context: Error messages now include step context (e.g., [Extraction - TypeDefinition] error...) for easier debugging

Derivation

  • HybridDerivation Base: All 13 modules use hybrid filtering combining pattern-based AND graph-based candidate selection
  • Edge-Aware Relationships: New Tier 1.5 derivation using CALLS, IMPORTS, USES edges with high confidence (0.90-0.95)
  • Pre-Generation Dedup: Fuzzy matching against existing elements before LLM calls
  • Business Layer: BusinessProcess detects orchestrator methods (3+ CALLS), BusinessEvent detects webhooks/signals, BusinessActor detects auth decorators
  • Relationship Consolidation: Refine step boosts confidence on multi-signal agreement, prunes low-confidence without corroboration
  • Self-Loop Prevention: Fixed self-referential relationships in graph_relationships with query filters and cleanup
  • Enrichment Cache: Aligned with LLM cache patterns, CLI control via --no-enrichment-cache

Adapters

  • Pydantic Structured Output: New schemas.py with Pydantic models for all extraction types. LLM manager auto-resolves JSON schemas to models, enforcing structure via PydanticAI
  • Rate Limiting: Adaptive throttling (auto-reduces RPM on 429s), circuit breaker pattern, Retry-After header respect, error classification, and model-specific rate limits via env vars
  • Graph Metadata: Element properties now include all graph metrics (kcore, articulation points, degree); propagated to relationships
  • Graph Labels: Split Neo4j labels into namespace (Graph/Model) and node type, enabling cleaner queries and consistent naming across extraction and derivation
  • Database Locking: Non-blocking during benchmarks/pipeline runs, versions used for isolation

Full Changelog: v0.6.8...v0.6.9