Releases: kinti/crawlward
Release list
v0.2.1 — verification correctness fixes
A review pass over v0.2.0 found three correctness bugs — two in --verify, the flagship feature. All fixed, each with a regression test (39 tests total, up from 27).
Fixed
- v4-mapped IPv6 addresses were reported as spoofed identities. nginx dual-stack (
ipv6only=off) logs IPv4 clients as::ffff:203.0.113.5; the IPv6 parser returned null for that form, so legitimate vendor traffic failed the IP-range check. The parser now canonicalizes v4-mapped addresses onto the IPv4 table — verified end-to-end against live OpenAI ranges. NAT64 (64:ff9b::…), zone indexes (fe80::1%eth0) and full-form dotted quads also parse now. - Zero-source-IP verification printed a misleading verdict (
0/0 — MIXED identity, inspect closely) when logs carry no client IPs. It now says explicitly: no source IPs in these logs — cannot verify (log remote_addr / remote_ip). Applebot-Extendedwas listed as a crawler. Apple's documentation states it "does not crawl webpages" — it is a robots.txt control token exactly likeGoogle-Extended, which this registry already documented as such. Moved tocontrolTokens(registry v4); the README's traps section now names both.
Also improved
- Undated events are excluded while
--since/--untilis active (previously they slipped through the filter);--helpdocuments that date-only values mean UTC midnight - Peak-hour buckets get their own (much higher) cap, so
peak/hno longer silently understates on logs longer than ~7 months - Unmatched bot-like UAs are version-normalized (
Scrapy/2.11andScrapy/2.12merge into one bucket) with a bounded map - gzip is detected by magic bytes, so compressed rotations work under any filename
- New Scope section in docs/crawlers.md explaining why
Googlebot/Bingbotare deliberately not classified as AI crawlers
Full changelog: v0.2.0...v0.2.1
v0.2.0 — identity verification & Apache support
The headline feature: --verify catches spoofed bot identities.
crawlward --verify /var/log/caddy/*.logA UA string is free text — anyone can claim to be GPTBot. Now crawlward
checks the source IPs of claimed-bot requests against the vendor's own
published IP prefix lists (OpenAI, Anthropic, Perplexity, Common Crawl and
Mistral publish them; IPv4 + IPv6 CIDR matching). The report now ends with a
verdict per bot: identity consistent, treat as spoofed, or mixed
(usually brand-new vendor ranges, not fraud).
Also in this release
- Apache combined log support — auto-detected per line, alongside Caddy
and nginx JSON; the client IP from combined logs feeds--verify - Peak requests/hour per bot — an aggressiveness signal you can act on
--since/--untildate filters and--csvoutput- Registry v3: each bot's verification URL is now data (
rangesfield),
with the annotated list in docs/crawlers.md - 27 tests; verification logic is tested with injected fetches, so CI stays
offline
Full changelog: v0.1.0...v0.2.0
v0.1.0 — first public release
First public release of crawlward, a watchdog for AI crawlers at your own origin — for everyone not on Cloudflare.
Highlights
- Access-log analyzer (zero dependencies, Node ≥ 18): classifies AI crawler traffic from Caddy and nginx JSON logs, plain or gzipped, with streaming reads that never load the whole file into memory.
- Rich report: per-bot and per-vendor aggregation, bandwidth, top paths, HTTP status mix, robots.txt awareness, and a bucket for bot-like user agents missing from the registry.
- Machine-readable registry with 39 crawler signatures, each verified against the vendor's official documentation on 2026-09-07. Unverifiable entries (e.g.
Bytespider) are clearly flagged; traps likeGoogle-Extended(a robots.txt-only token with no HTTP user-agent) are documented to prevent classic reporting mistakes. --jsonoutput for pipelines and dashboards.- Case-insensitive token matching (Meta's wire UA strings are lowercase).
- 14 tests (
node:test), CI on Node 18/20/22.
Install
git clone https://github.com/kinti/crawlward.git
node crawlward/src/analyze.mjs /var/log/caddy/*.logSee the README for Caddy/nginx log setup.