Skip to content

Releases: danishxsethi/blazecrawl

BlazeCrawl v0.1.2

Choose a tag to compare

@danishxsethi danishxsethi released this 22 Sep 22:06
8ac9932

Fixed

  • Corrected BlazeCrawl Core runtime version reporting: blazecrawl_core.__version__, blazecrawl --version, and /health now report the actual release version (the published blazecrawl-core==0.1.1 wheel incorrectly reported 0.1.0 at runtime).
  • Aligned the Node SDK and MCP package-lock.json versions with their package versions.
  • Strengthened release-integrity validation (manifests, both Node lockfiles, and built-artifact runtime versions) to prevent future version drift.

Distribution

Available via:

Compatibility

No intentional scrape/map/crawl API behavior change.

BlazeCrawl v0.1.0

Choose a tag to compare

@danishxsethi danishxsethi released this 22 Sep 02:27

BlazeCrawl v0.1.0

BlazeCrawl is a security-first, self-hostable web extraction toolkit for developers and AI systems.

What is BlazeCrawl?

BlazeCrawl Core is the open-source engine for turning web pages into clean, LLM-ready Markdown and structured data. It's built for people who want to run their own scraping infrastructure without handing their URLs, traffic, or credentials to a third party — and without giving up the outbound-request security that most self-hosted scrapers skip.

Included in v0.1.0

  • Three core endpoints: /v1/scrape, /v1/map, /v1/crawl
  • Security-first egress: SSRF validation, DNS-rebinding resistance, private/loopback/cloud-metadata blocking, redirect re-validation, browser request interception
  • Hybrid rendering: Fast static fetch with Playwright browser fallback for JS-heavy pages
  • Clean Markdown extraction: readability → trafilatura → BeautifulSoup
  • robots.txt compliance: Same-origin BFS crawler with robots enforcement
  • SDKs: Python SDK, Node SDK, CLI
  • MCP servers: Python and Node implementations
  • Docker self-hosting: Zero-config docker-compose setup
  • Local API key auth: Auto-generated, persisted, secure permissions

Security Model

BlazeCrawl treats every user-supplied URL as hostile:

  • SSRF-aware destination validation: Resolves once, validates all returned IPs against private/reserved denylist
  • Browser egress proxy: All browser subresource requests go through validated proxy
  • DNS-rebinding protection: Connection pinned to validated IP
  • Security regression suites: WP1B (HTTP egress), WP2C (HTTPS egress), WP3 (browser subresource isolation)

See docs/SECURITY_MODEL.md for complete details.

Quickstart

git clone https://github.com/danishxsethi/blazecrawl.git
cd blazecrawl
docker compose up --build

# Get API key from logs
docker compose logs api | grep "first run"

# Scrape a page
curl -X POST http://localhost:8000/v1/scrape \
  -H "Authorization: Bearer blz_local_xxx" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://example.com"}'

v0.1 Limitations

  • Single-node architecture: No distributed crawl coordination
  • No managed anti-bot: Deliberately excluded (commercial feature)
  • No managed proxy network: Deliberately excluded (commercial feature)
  • WebSockets blocked: Browser security policy prevents WebSocket connections
  • Workers blocked: Web Workers, SharedWorkers, ServiceWorkers blocked
  • No hosted SLA: Community support only
  • Platform: Tested on Linux x86_64 only

See docs/OSS_VS_CLOUD.md for what's deliberately excluded.

Contributing

We welcome contributions! See:

License

Apache-2.0. See LICENSE and THIRD_PARTY_LICENSES.md.