Releases: danishxsethi/blazecrawl
Release list
BlazeCrawl v0.1.2
Fixed
- Corrected BlazeCrawl Core runtime version reporting:
blazecrawl_core.__version__,blazecrawl --version, and/healthnow report the actual release version (the publishedblazecrawl-core==0.1.1wheel incorrectly reported0.1.0at runtime). - Aligned the Node SDK and MCP
package-lock.jsonversions with their package versions. - Strengthened release-integrity validation (manifests, both Node lockfiles, and built-artifact runtime versions) to prevent future version drift.
Distribution
Available via:
- PyPI
blazecrawl-core— server / CLI - PyPI
blazecrawl— Python SDK - PyPI
blazecrawl-mcp— Python MCP server - npm
@blazecrawl/sdk— Node SDK - npm
@blazecrawl/mcp— Node MCP server - GHCR
ghcr.io/danishxsethi/blazecrawl:0.1.2— Linux x86_64 image
Compatibility
No intentional scrape/map/crawl API behavior change.
BlazeCrawl v0.1.0
BlazeCrawl v0.1.0
BlazeCrawl is a security-first, self-hostable web extraction toolkit for developers and AI systems.
What is BlazeCrawl?
BlazeCrawl Core is the open-source engine for turning web pages into clean, LLM-ready Markdown and structured data. It's built for people who want to run their own scraping infrastructure without handing their URLs, traffic, or credentials to a third party — and without giving up the outbound-request security that most self-hosted scrapers skip.
Included in v0.1.0
- Three core endpoints:
/v1/scrape,/v1/map,/v1/crawl - Security-first egress: SSRF validation, DNS-rebinding resistance, private/loopback/cloud-metadata blocking, redirect re-validation, browser request interception
- Hybrid rendering: Fast static fetch with Playwright browser fallback for JS-heavy pages
- Clean Markdown extraction: readability → trafilatura → BeautifulSoup
- robots.txt compliance: Same-origin BFS crawler with robots enforcement
- SDKs: Python SDK, Node SDK, CLI
- MCP servers: Python and Node implementations
- Docker self-hosting: Zero-config docker-compose setup
- Local API key auth: Auto-generated, persisted, secure permissions
Security Model
BlazeCrawl treats every user-supplied URL as hostile:
- SSRF-aware destination validation: Resolves once, validates all returned IPs against private/reserved denylist
- Browser egress proxy: All browser subresource requests go through validated proxy
- DNS-rebinding protection: Connection pinned to validated IP
- Security regression suites: WP1B (HTTP egress), WP2C (HTTPS egress), WP3 (browser subresource isolation)
See docs/SECURITY_MODEL.md for complete details.
Quickstart
git clone https://github.com/danishxsethi/blazecrawl.git
cd blazecrawl
docker compose up --build
# Get API key from logs
docker compose logs api | grep "first run"
# Scrape a page
curl -X POST http://localhost:8000/v1/scrape \
-H "Authorization: Bearer blz_local_xxx" \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com"}'v0.1 Limitations
- Single-node architecture: No distributed crawl coordination
- No managed anti-bot: Deliberately excluded (commercial feature)
- No managed proxy network: Deliberately excluded (commercial feature)
- WebSockets blocked: Browser security policy prevents WebSocket connections
- Workers blocked: Web Workers, SharedWorkers, ServiceWorkers blocked
- No hosted SLA: Community support only
- Platform: Tested on Linux x86_64 only
See docs/OSS_VS_CLOUD.md for what's deliberately excluded.
Contributing
We welcome contributions! See:
- CONTRIBUTING.md for development setup
- Issues for good first issues
- SECURITY.md for reporting vulnerabilities
License
Apache-2.0. See LICENSE and THIRD_PARTY_LICENSES.md.