See what crawlers cost your origin before you decide what to block.
Access logs show footprints. They rarely show the bill.
CrawlLedger reads the Nginx or Caddy logs you already have and turns them into something you can reason about: which claimed crawlers arrived, where they wandered, how much origin work followed, whether they fell into a crawl trap, and what a deny or rate-limit policy would have changed.
Everything stays local. The result is a self-contained HTML report, a machine-readable JSON report and—only after a separate simulation—reviewable server-configuration fragments. No telemetry. No daemon. Nothing in the request path, and no command that quietly rewrites a live server.
CI checks formatting, runs go vet, and tests on Go 1.26.5 across Ubuntu,
macOS and Windows. It also runs the race detector, eight fuzz smoke targets,
validates generated Nginx and Caddy fragments, exercises analysis without a
network namespace, cross-builds six pure-Go binaries, and checks schemas, docs
and module integrity.
This is the actual HTML report from the synthetic fixture included in the repository. No production traffic is shown.
- Claimed search, AI/LLM, monitoring and other crawler traffic, broken down without pretending a User-Agent is proof of identity.
- Origin load by normalized route: response bytes, duration, upstream time and cache state whenever the source log actually contains those fields.
- Deterministic findings for crawl traps, expensive 404s, query-space
explosion, cache busting, security probes and
robots.txtviolations. - Historical policy simulation with ordered, first-match rules.
- Nginx or Caddy configuration drafts. Drafts—not auto-applied changes.
- Privacy-conscious storage; raw log lines are never written to the workspace.
There is one hard boundary. A request labelled Googlebot, GPTBot, or
anything else has merely claimed that name through its User-Agent header.
CrawlLedger does not DNS-verify it. The label may be spoofed.
One prerequisite: Go 1.26.5.
make build
./bin/crawlledger analyze \
--input testdata/logs/nginx-combined/valid.log \
--format nginx-combined \
--output ./demo-auditOpen demo-audit/report.html in a browser. The fixture is tiny—three synthetic
requests—but it still demonstrates crawler classification, route
normalization, missing-metric disclosure and a security-probe finding.
Output paths must be new. Deliberately. CrawlLedger refuses to replace an existing workspace, simulation, evidence bundle, or rendered configuration.
There is no public v1 tag yet, so build from this checkout for now. Tagged releases will carry binaries and checksums on GitHub Releases.
Nginx combined log? Point CrawlLedger at it:
crawlledger analyze \
--input /var/log/nginx/access.log \
--format nginx-combined \
--output ./crawl-auditCaddy JSON is just as direct:
crawlledger analyze \
--input /var/log/caddy/access.log \
--format caddy-json \
--output ./crawl-auditRepeat --input to process several files in order. Plain text and gzip are
detected from their contents, not from a hopeful filename suffix. Use
--input - once to read standard input.
One run, one explicit format. CrawlLedger does not guess; a firm parse error is safer than a polished report built from the wrong grammar.
| Format | Use it for | Metric coverage |
|---|---|---|
nginx-combined |
The standard Nginx combined format | Requests, status, bytes, referer and User-Agent |
nginx-json |
The CrawlLedger Nginx JSON format below | Adds request time, upstream time and cache state |
caddy-json |
Standard Caddy JSON access logs | Adds request duration; upstream and cache metrics are unavailable |
crawlledger-json |
A bundle created by crawlledger sanitize |
Preserves the metrics present in the source bundle |
Recommended Nginx JSON log format
Add this inside the Nginx http block:
log_format crawlledger escape=json
'{"timestamp":"$time_iso8601",'
'"remote_addr":"$remote_addr",'
'"method":"$request_method",'
'"uri":"$request_uri",'
'"status":"$status",'
'"bytes_sent":"$body_bytes_sent",'
'"request_time":"$request_time",'
'"upstream_response_time":"$upstream_response_time",'
'"upstream_cache_status":"$upstream_cache_status",'
'"user_agent":"$http_user_agent",'
'"referer":"$http_referer"}';
access_log /var/log/nginx/access.crawlledger.json crawlledger;Then analyze it:
crawlledger analyze \
--input /var/log/nginx/access.crawlledger.json \
--format nginx-json \
--output ./crawl-audit| File | Purpose |
|---|---|
report.html |
Local, self-contained report for a browser or print-to-PDF |
report.json |
Versioned report for scripts and further analysis |
analysis.sqlite |
Normalized aggregates and bounded parse diagnostics |
manifest.json |
Completion marker, artifact sizes and SHA-256 hashes |
No manifest.json? Treat the workspace as incomplete.
Analysis tells you what happened. Simulation asks the more dangerous question: what would a policy have done?
Policies are strict, versioned JSON. Rules run from top to bottom; the first
match wins. Small examples live in
testdata/policies.
crawlledger policy simulate \
--workspace ./crawl-audit \
--policy ./testdata/policies/safe.json \
--output ./simulation.jsonRead the totals. Then the rule impacts, coverage and assumptions. Most of all,
read every item in risks.
Some risks demand an explicit acknowledgement before rendering:
crawlledger policy render \
--workspace ./crawl-audit \
--policy ./testdata/policies/safe.json \
--simulation ./simulation.json \
--target nginx \
--output ./rendered-nginx \
--ack claimed-allow-bypass:allow-googlebotThat acknowledgement belongs to the bundled example. Use the exact IDs from
your own simulation.json; unknown IDs and non-acknowledgeable blockers are
rejected.
Nginx drafts support allow, deny and fixed rate profiles. Caddy drafts
support allow and deny. Cache rules are analytical upper bounds, so a
policy containing one remains simulation-only and cannot be rendered.
Draft means draft. CrawlLedger never runs Nginx, Caddy, Docker, systemd, SSH, or a shell.
For Nginx, include crawlledger-http.conf inside http {} and
crawlledger-server.conf inside the audited server {}, before the content
handler. Back up the current configuration. Then validate:
nginx -tFor Caddy, import Caddyfile.crawlledger before reverse_proxy, then run:
caddy adapt --config Caddyfile --adapter caddyfile --validate
caddy validate --config CaddyfileReload manually and watch responses as well as access logs. If validation or traffic turns strange, restore the backup and reload.
A raw access log is useful evidence wrapped around sensitive data. Do not pass it around casually. Create a pseudonymized bundle instead:
crawlledger sanitize \
--input /var/log/nginx/access.log \
--format nginx-combined \
--output ./crawl-evidence \
--key-file /secure/crawlledger.keyIf the key file does not exist, CrawlLedger creates a private-permission,
32-byte key. Keep it outside both bundle and workspace. Reusing the key makes
pseudonyms comparable across runs; omitting --key-file creates an ephemeral
key that is not retained.
Before any record reaches disk, CrawlLedger:
- replaces client IPs and User-Agent values with domain-separated HMAC pseudonyms;
- removes query values while keeping safe query-key names;
- normalizes token-like path segments;
- cuts referers down to hostnames;
- discards cookies and raw log lines.
Safer is not anonymous. The bundle still holds timestamps, normalized routes, crawler claims, referer hosts and stable pseudonyms; those can be correlated or re-identified. Treat bundles, workspaces, simulations, keys and generated configuration as sensitive files.
- Claimed crawler identity is spoofable. CrawlLedger audits traffic; it does not authenticate bots.
- Analysis is batch-based. There is no live protection.
- Combined Nginx and standard Caddy logs cannot expose every origin or cache metric. Reports show the holes instead of filling them with guesses.
- Cost figures are proportional allocations driven by your configuration, neither invoices nor promised savings.
- Cache simulation cannot see application semantics such as
Authorization,Set-Cookie,Cache-Control, orVary. - History is not prophecy: a simulation cannot know how a crawler will react after the policy changes.
CrawlLedger is released under the MIT License. Third-party attributions live in NOTICE.
The everyday local gate is short:
make checkThat checks formatting, runs go vet, tests the code and builds the project.
Before a release—or after changing parsing, normalization, policy evaluation,
or rendering—go further:
make test-race
make fuzz-smokeFixtures use synthetic documentation addresses and .example domains. Keep
production logs, customer workspaces, credentials and real HMAC keys out of
the repository.
Read CONTRIBUTING.md before touching schemas, migrations, or the embedded crawler catalog. Security reports take a different route: follow SECURITY.md, and never place sensitive details in a public issue.
This project was developed with AI assistance and is maintained by the author.
