Skip to content

TextSpitter v1.0.0

Latest

Choose a tag to compare

@fsecada01 fsecada01 released this 18 Feb 05:21
· 10 commits to main since this release

Production Release 🚀

TextSpitter v1.0.0 is now available on PyPI!

What's New

Core Features

  • 📄 Multi-format extraction: PDF (PyMuPDF + PyPDF fallback), DOCX, TXT, CSV, and 50+ source code file types
  • 🔌 Stream-first API: Supports file paths, BytesIO, SpooledTemporaryFile, and raw bytes
  • 🛠️ Optional logging: Built-in loguru support with automatic stdlib logging fallback
  • 🖥️ CLI tool: textspitter command for quick one-off extractions
  • 📚 Professional documentation: Landing page, quick-start, tutorials, use cases, and recipes

Quality Assurance

  • 🧪 ~80 pytest tests with 89%+ code coverage
  • ✨ Full type hints with PEP 561 marker
  • 🔒 Pre-commit hooks (ruff, black, isort, ty)
  • 🚀 Automated CI/CD on Python 3.12–3.14
  • 📖 Auto-published docs to GitHub Pages

Installation

pip install textspitter

# With optional loguru logging
pip install "textspitter[logging]"

Quick Example

from TextSpitter import TextSpitter

text = TextSpitter(filename="report.pdf")
print(text[:200])

Documentation

Key Improvements in v1.0.0

✅ Fixed CI/CD type checking for Python 3.12–3.14
✅ Added system dependencies for lxml build on Python 3.14
✅ Replicated SQLModel-CRUD-Utilities documentation approach
✅ Created professional landing page with feature cards
✅ Complete documentation pages with proper navigation
✅ Fixed code block whitespace formatting
✅ All pre-commit hooks passing

Thanks for using TextSpitter! 🎉