Two open-source Crawl4AI plugins for AI coding agents - Claude Code + pi (with Trafilatura, ~60% fewer tokens) #2316
romek-rozen
started this conversation in
Show and tell
Replies: 1 comment
|
Hey @romek-rozen 😊 Thank you for this. It's a great contribution! 🙌 We like the idea. It's a smart way to cut the boilerplate. It's also great to see both Claude Code and pi covered. We're checking out both plugins. If everything looks good, we'd love to bring this in. We'll get back to you soon with our thoughts. Thanks again for building this and sharing it with the community. It means a lot 💜 |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hi everyone 👋
I built with Codex and Cloud two small open-source plugins that bring Crawl4AI into AI coding agents, and I wanted to share them with the community.
The problem
When an agent reads a web page, most of what it gets is boilerplate: navigation, menus, cookie banners, leftover scripts. All of that eats tokens and fills up the context window.
The idea
Crawl4AI renders the page, then Trafilatura strips the boilerplate locally, and optional BM25 filtering keeps only the parts that match the query. The agent gets a path to a compact artifact instead of the whole page dumped into the conversation.
Real example: the Claude Code features overview page (~500 KB of HTML):
markdown-fit(cl100k_base tokenizer, so the numbers are indicative. Results vary by site.)
1. cc-crawl4ai: plugin for Claude Code
https://github.com/romek-rozen/cc-crawl4ai
2. pi-crawl4ai: package for the pi coding agent
https://github.com/romek-rozen/pi-crawl4ai
crawl4aitool, plus/crawl4ai-install,/crawl4ai-statusand/crawl4ai-testcommandsBoth are MIT licensed. Feedback, issues, and ideas are very welcome. A big thank-you to @unclecode and the Crawl4AI team for such a great library 🙏
All reactions