A review pass over v0.2.0 found three correctness bugs — two in --verify, the flagship feature. All fixed, each with a regression test (39 tests total, up from 27).
Fixed
- v4-mapped IPv6 addresses were reported as spoofed identities. nginx dual-stack (
ipv6only=off) logs IPv4 clients as::ffff:203.0.113.5; the IPv6 parser returned null for that form, so legitimate vendor traffic failed the IP-range check. The parser now canonicalizes v4-mapped addresses onto the IPv4 table — verified end-to-end against live OpenAI ranges. NAT64 (64:ff9b::…), zone indexes (fe80::1%eth0) and full-form dotted quads also parse now. - Zero-source-IP verification printed a misleading verdict (
0/0 — MIXED identity, inspect closely) when logs carry no client IPs. It now says explicitly: no source IPs in these logs — cannot verify (log remote_addr / remote_ip). Applebot-Extendedwas listed as a crawler. Apple's documentation states it "does not crawl webpages" — it is a robots.txt control token exactly likeGoogle-Extended, which this registry already documented as such. Moved tocontrolTokens(registry v4); the README's traps section now names both.
Also improved
- Undated events are excluded while
--since/--untilis active (previously they slipped through the filter);--helpdocuments that date-only values mean UTC midnight - Peak-hour buckets get their own (much higher) cap, so
peak/hno longer silently understates on logs longer than ~7 months - Unmatched bot-like UAs are version-normalized (
Scrapy/2.11andScrapy/2.12merge into one bucket) with a bounded map - gzip is detected by magic bytes, so compressed rotations work under any filename
- New Scope section in docs/crawlers.md explaining why
Googlebot/Bingbotare deliberately not classified as AI crawlers
Full changelog: v0.2.0...v0.2.1