Skip to content

v0.2.1 — verification correctness fixes

Latest

Choose a tag to compare

@kinti kinti released this 13 Sep 23:11
· 7 commits to main since this release

A review pass over v0.2.0 found three correctness bugs — two in --verify, the flagship feature. All fixed, each with a regression test (39 tests total, up from 27).

Fixed

  • v4-mapped IPv6 addresses were reported as spoofed identities. nginx dual-stack (ipv6only=off) logs IPv4 clients as ::ffff:203.0.113.5; the IPv6 parser returned null for that form, so legitimate vendor traffic failed the IP-range check. The parser now canonicalizes v4-mapped addresses onto the IPv4 table — verified end-to-end against live OpenAI ranges. NAT64 (64:ff9b::…), zone indexes (fe80::1%eth0) and full-form dotted quads also parse now.
  • Zero-source-IP verification printed a misleading verdict (0/0 — MIXED identity, inspect closely) when logs carry no client IPs. It now says explicitly: no source IPs in these logs — cannot verify (log remote_addr / remote_ip).
  • Applebot-Extended was listed as a crawler. Apple's documentation states it "does not crawl webpages" — it is a robots.txt control token exactly like Google-Extended, which this registry already documented as such. Moved to controlTokens (registry v4); the README's traps section now names both.

Also improved

  • Undated events are excluded while --since/--until is active (previously they slipped through the filter); --help documents that date-only values mean UTC midnight
  • Peak-hour buckets get their own (much higher) cap, so peak/h no longer silently understates on logs longer than ~7 months
  • Unmatched bot-like UAs are version-normalized (Scrapy/2.11 and Scrapy/2.12 merge into one bucket) with a bounded map
  • gzip is detected by magic bytes, so compressed rotations work under any filename
  • New Scope section in docs/crawlers.md explaining why Googlebot/Bingbot are deliberately not classified as AI crawlers

Full changelog: v0.2.0...v0.2.1