v0.7.0
Normalized Semantic Chunker v0.7.0
✨ New Features
JSON File Support: Processing of JSON files with format {"chunks": [{"text": "..."}, ...]}
Dynamic Memory Management: Smart worker allocation based on system resources
Verbosity Controls: Configurable logging for debugging and production
Configurable Parameters: Control via environment variables
File Validation: Input size and format checks
🚀 Performance Improvements
Batch Processing: Prevents OOM errors for large documents (>20K sentences)
Model Caching: Cache system with automatic expiration (1h default)
Adaptive Workers: Scalability based on document size
Memory Cleanup: Optimized GPU memory management
Adaptive Step Size: Optimization based on document size
🛡️ Robustness
Error Handling: Smart fallback mechanisms for tiktoken errors
Automatic Recovery: Recovery mechanisms from processing failures
Improved Logging: Detailed and configurable logging system
Input Validation: Comprehensive checks on file size, format, and content
📈 Improvement Metrics
⬇️ Memory Usage: -30–40% for large documents
⚡ Speed: +15–25% for documents >10K sentences
🛠️ Reliability: +95% reduction in processing errors
🔧 Configurability: Full control via environment variables