v0.5.0 - Automatic Text Chunking
What's New
This release adds automatic text chunking with aggregate embeddings for handling documents that exceed embedding model token limits.
Features
-
Automatic Text Chunking: New chunker package with configurable strategies
FixedOverlapChunker: splits text with configurable size and overlap- Uses tiktoken for accurate token counting (cl100k_base encoding)
- Default: 512 token chunks with 50 token overlap, 8191 max tokens
-
Batch Embedding Support: Efficient multi-text embedding
- New
BatchEmbeddingProviderinterface - Implemented in OpenAI provider (up to 2048 texts per request)
- Automatic fallback to individual embeddings if batch not supported
- New
-
Aggregate Embedding Approach
- Chunks text when exceeding token limits
- Embeds all chunks using batch API for performance
- Averages chunk embeddings into single aggregate embedding
- Stores as single entry (no derived keys needed)
-
Improved Precision: Converted from float32 to float64
- Uses OpenAI's native float64 format
- Updated all similarity functions and backends
- Better precision for similarity calculations
Configuration
- Chunking enabled by default with sensible defaults
WithChunking()to customize chunk size/overlap/strategyWithoutChunking()to disable if not needed- Zero-config works out of the box
Testing
- Comprehensive chunker package tests
- Cache integration tests for aggregate embeddings
- All existing tests pass
Full Changelog: v0.4.0...v0.5.0