Repository navigation
Extractium v0.3
Extractium™
Description
Extractium™ turns your organization's scattered public documentation into one searchable knowledge base. Point it at your website, knowledge base portal, GitHub repositories, YouTube channel, library repository, or a folder of files, and it gathers the content, prepares it for both keyword and meaning-based search, and writes it out in several formats. You can then use that knowledge base in a website search box, in your own scripts, or with the AI assistant of your choice, without depending on any one AI provider.
Unlike a vector database, Extractium™ needs no server, no database, and no API to run. Every output is a static file that you can host anywhere, including GitHub Pages, and the same build feeds all of them at once. Sources and outputs are plug-ins, so you can add your own if the built-in ones do not cover your needs.
Learn more at: github.com/DepressionCenter/extractium.
What's Changed
- Raise the version to 0.2 and read it from one place by Gabriel Mongefranco (@gabrielmongefranco) in #45
- Stop crawling the narrowed views of a portal question listing by Gabriel Mongefranco (@gabrielmongefranco) in #46
- Skip an account's housekeeping repositories and accept their names by Gabriel Mongefranco (@gabrielmongefranco) in #47
- Read UTF-16 repository files and say what max_pages counts by Gabriel Mongefranco (@gabrielmongefranco) in #48
- Let a web source add exclude patterns without replacing the defaults by Gabriel Mongefranco (@gabrielmongefranco) in #49
- Carry the code parsers in the lock file and read Windows batch files by Gabriel Mongefranco (@gabrielmongefranco) in #51
- Index a long file as a compact record, and keep a page's text past the parse ceiling by Gabriel Mongefranco (@gabrielmongefranco) in #52
- Let the automatic transport post, and keep program folders out of a crawl by Gabriel Mongefranco (@gabrielmongefranco) in #53
- Run sources at the same time and keep several page fetches in flight by Gabriel Mongefranco (@gabrielmongefranco) in #60
- Choose a standard Python for the build environment by Gabriel Mongefranco (@gabrielmongefranco) in #61
- Build on a free-threaded Python, and check the Python the script was given by Gabriel Mongefranco (@gabrielmongefranco) in #62
- Read shared Google files and Word, OpenDocument, and RTF documents by Gabriel Mongefranco (@gabrielmongefranco) in #64
- Add leaf patterns for single pages on other hosts by Gabriel Mongefranco (@gabrielmongefranco) in #65
- Carry the caption library in the lock file by Gabriel Mongefranco (@gabrielmongefranco) in #67
- Parse Go and Rust with their published grammars by Gabriel Mongefranco (@gabrielmongefranco) in #68
- Read PDF files with pypdf in a killable child process by Gabriel Mongefranco (@gabrielmongefranco) in #69
- Read portal attachments, carry the enrichment fields, and fix the caption session by Gabriel Mongefranco (@gabrielmongefranco) in #70
- Apply max_pages to every source, in the unit each one reads by Gabriel Mongefranco (@gabrielmongefranco) in #71
- Record a redirected page where it landed, recheck pypdf security advisories, and untrack the lint reports by Gabriel Mongefranco (@gabrielmongefranco) in #72
- Name every section with keywords and every page with tags using YAKE by Gabriel Mongefranco (@gabrielmongefranco) in #73
- Read PowerPoint and OpenDocument presentations one section per slide by Gabriel Mongefranco (@gabrielmongefranco) in #74
Full Changelog: v0.2...v0.3
Copyright © 2026 The Regents of the University of Michigan
