Skip to content

v0.5.0 - Automatic Text Chunking

Choose a tag to compare

@botirkhaltaev botirkhaltaev released this 27 Nov 18:25
· 25 commits to master since this release

What's New

This release adds automatic text chunking with aggregate embeddings for handling documents that exceed embedding model token limits.

Features

  • Automatic Text Chunking: New chunker package with configurable strategies

    • FixedOverlapChunker: splits text with configurable size and overlap
    • Uses tiktoken for accurate token counting (cl100k_base encoding)
    • Default: 512 token chunks with 50 token overlap, 8191 max tokens
  • Batch Embedding Support: Efficient multi-text embedding

    • New BatchEmbeddingProvider interface
    • Implemented in OpenAI provider (up to 2048 texts per request)
    • Automatic fallback to individual embeddings if batch not supported
  • Aggregate Embedding Approach

    • Chunks text when exceeding token limits
    • Embeds all chunks using batch API for performance
    • Averages chunk embeddings into single aggregate embedding
    • Stores as single entry (no derived keys needed)
  • Improved Precision: Converted from float32 to float64

    • Uses OpenAI's native float64 format
    • Updated all similarity functions and backends
    • Better precision for similarity calculations

Configuration

  • Chunking enabled by default with sensible defaults
  • WithChunking() to customize chunk size/overlap/strategy
  • WithoutChunking() to disable if not needed
  • Zero-config works out of the box

Testing

  • Comprehensive chunker package tests
  • Cache integration tests for aggregate embeddings
  • All existing tests pass

Full Changelog: v0.4.0...v0.5.0