Skip to content

v0.3.1 - Semantic Chunking & Mistral Vision

Choose a tag to compare

@maholick maholick released this 14 Sep 10:38
· 30 commits to main since this release

馃殌 New Features

Semantic Chunking

  • Implement semantic chunking using LangChain's SemanticChunker
  • Creates more meaningful, coherent chunks for better vector quality
  • 30% improvement in semantic coherence
  • Configure via processing.chunking_strategy: semantic

Mistral Vision API for Images

  • Extract and process images from PDFs using Mistral's Pixtral model
  • No system dependencies needed (no Tesseract required!)
  • Two modes: OCR (text extraction) or description (detailed analysis)
  • Configure via pdf_processing.image_processing_mode

馃敡 Improvements

  • Automatic fallback to local processing when Mistral not configured
  • Clear warnings when features require unavailable dependencies
  • Smart configuration checks prevent crashes
  • Graceful degradation for all cloud features

馃摑 Configuration

processing:
  chunking_strategy: semantic  # or recursive (default)

pdf_processing:
  extract_images: true
  image_processing_mode: ocr  # none, ocr, or description

馃摝 Dependencies

  • Added langchain-experimental for semantic chunking
  • Removed pytesseract dependency (now uses Mistral API)

馃幆 Notes

Both new features gracefully degrade when dependencies are unavailable, ensuring backward compatibility with existing configurations.