Repository navigation
geo-score 1.5.0
A release about reading sites the way the rubric says, and about failing CI for the right reason. The two
crawler gates now read only the ten retrieval crawlers, choose each crawler's robots.txt group as RFC 9309
does and retry a request that got no answer before they count it; a bot-challenge page is no longer scored
as a page that needs JavaScript; every page is read in its own language; and a figure counts as sourced
only when its source sits beside it. These fixes move scores, and each one below names its check and its
direction. For CI there are --fail-on-gate and --assert (a gate at zero caps a score at 40, and a site
that would otherwise score 40 or more reads exactly 40, which --fail-under 40 lets through),
geo-score diff and --baseline to compare two reports check by check, and a shields.io endpoint badge.
A site that cannot be scored fails fast and says why in one line, in JSON under --json, and a scoring
run can no longer hang. Around the CLI: a schema and typo warnings for the watch config, installs from the
release wheel without git, a gate that pins its tools, CI examples beyond GitHub, forms for disputing a
leaderboard row, and a benchmark that can be re-measured and taken apart. The leaderboard itself is still
the 1.3.0 run of 2026-09-27.
Upgrading from 1.4.0. Level-1 scores move only where 1.4.0 read a site wrong, and every such fix is
under Fixed with its check and direction: the gates (g.robots, g.reachable, and g.ssr for sites
behind a bot challenge, with Japanese or Korean pages, or with English pages that carry more than 20 Han
characters), the checks that read page text (p2.answer-passages, p2.question-intent,
p2.sourced-stats, p2.named-author, p4.answer-shape, p4.cn-engines and p3.knowledge-graph), and
the checks whose file, logo or crawler request is now retried (p1.llms-txt, b.llms-full, b.ai-txt,
p1.organization and p4.cn-engines). Three invocations exit differently: a site that serves only
bot-challenge pages exits 2 instead of being scored; an unexpected error exits 2 in one line instead of
a traceback that exited 1; and --sample below 1 exits 2, where 1.4.0 crashed on 0 and scored fewer
pages on a negative value. diff is now a subcommand, so a host named diff is written
https://diff. audience_language can hold languages other than en and zh-CN, and the report gains
only the optional elapsed_s.
Added
- A glossary and a Chinese CLI reference.
guide/glossary.md(andguide/glossary.zh-CN.md) defines the 49
terms the README, the reports and the guides use, each with where it appears and which page goes deeper;
cli/README.zh-CN.mdis the CLI reference in Simplified Chinese, with a language switch at the top of both. --fail-on-gateand--assert EXPR, in the CLI and as the Action inputsfail-on-gateand
assert. A gate at zero caps the score at 40: a site that would otherwise score 40 or more reads
exactly 40 and passes--fail-under 40(fail-under: 40).--fail-on-gatefails any capped site: it
exits 1 when anyg.*check scores 0, and names the gate and its evidence; it is the same as
--assert 'g.*>0'.--assertholds one check, or an fnmatch glob of checks, to a value, e.g.
g.robots=maxorp1.llms-txt>=4, and is repeatable. A check the run could not observe is skipped,
never failed. A malformed expression, or a glob that names no check, exits 2.
Both are opt-in and change no score, and their results go to stderr, never into the report.geo-score diff BEFORE AFTER [--format text|md|json] [--fail-on-drop]compares two level-1 reports
check by check. Each check shows its points before and after, its max, and a state:up,down,
same,became unobservable,became observable, orunobservablewhen neither report observed it.
The comparison also shows the score, band and gate cap before and after, names each gate that reached
or left 0, and lists the pages added to or removed from the sample. When the two samples differ, it
says so, reads the totals as ±5 (benchmark/REPRODUCIBILITY.md) and namesg.reachableas the least
stable check. A check that became unobservable is never shown as a drop to 0.--format jsonfollows
the newschema/diff.v1.json. Two reports of different rubric versions exit 2, and reports from
different releases, or of two different targets, are compared with a warning on stderr.diffneeds
onlygeo_score.py.--baseline REPORTon a scoring run scores the baseline's pages again (as--urls-fromdoes) and
prints the same comparison after the report. It cannot be combined with--urlsor--urls-from.
Under--json, stdout carries only the report and the comparison is not printed: write the report
with--json-outand rungeo-score diffon the two files to see it.--fail-on-dropstill exits 4
and names each drop on stderr.--fail-on-drop(with--baseline, or ondiff) exits 4 when a check lost points, a gate reached
0 or the band fell, and prints one stderr line per reason. It is opt-in, aswatch diff --fail-on-drop
is, so no existing invocation changes its exit code. A failed--fail-under,--fail-on-gateor
--assertstill exits 1.- Action: a
baselineinput,deltaanddropped-checksoutputs, and a "Changes since baseline"
section in the job summary, with the check ids read from the baseline escaped. The baseline may be the
same file asjson-out. - A badge that stays current.
--badge-json FILEwrites the badge as shields.io
endpoint JSON:
{"schemaVersion": 1, "label": "AIV readiness", "message": "71/100 Solid", "color": "#2f8f52", "cacheSeconds": 86400}.
The SVG and the JSON take their message and colour from one helper, so the two cannot disagree. Host
the file anywhere andhttps://img.shields.io/endpoint?url=<its URL>draws the badge. The Action gets
abadgeinput: a path ending in.jsongets the endpoint JSON, and any other path gets the SVG. A
badge-pathoutput says where the file went. - A JSON error instead of nothing. Under
--json, every exit 2, usage errors included, prints
{"error": {"kind", "message", "target"}, "tool": "geo-score-cli/X.Y.Z"}on stdout, where stdout used
to be empty;diff --format jsondoes the same. The newschema/error.v1.jsonnames the kinds:dns,
refused,tls,timeout,droppedandno_responsewhen nothing answered;http_<code>,
no_content,slowandbot_challengewhen the site answered;deadlinewhen the run ran out of
time;not_publicunder the MCP server's address guard; andusage,rubric_mismatch(reports of two
rubric versions given todiffor--baseline),errorandinternal. A minor release may add
kinds; read an unknown one aserror. --timeout SECONDS(default 15, as before; up to 300): how long each request may wait for the
server before it counts as unanswered.elapsed_sin the JSON report (optional,schema/report.v2.json): how many seconds the run took
on the machine that ran it. The footer prints it too (· took 23.4 s). It describes the run, not the
site, anddiffignores it.- A schema for
geo-score-watch.json, and a warning for every key nothing reads.
schema/watch-config.v1.json(JSON Schema draft 2020-12) describes the file, andwatch initnow
writes"$schema"as its first key, so an editor checks and completes it. A mistyped key used to
vanish without a word:budgetforbudget_usdmeant no budget, andmax_callformax_callsthe
default cap of 120.watch checknames every unknown key at every level (top,brand,
competitors[i],engines[i],prices.<id>), with the nearest known key when one is close, as in
unknown key 'budget' (did you mean 'budget_usd'?); it is ignored, and flags apricesentry that
names no engine.watch runand the other commands that read the config print the same warnings on
stderr; the MCPstatustool lists them, andrunshows them with its plan and after a saved run. A
key named like a credential (api_key,OPENAI_KEY,api_token,client_secret,credentials…)
is pointed at environment variables instead, and its value is never printed. Warnings only: no exit
code changes. - CI beyond GitHub.
examples/ci/gitlab-ci.ymlis a GitLab job onpython:3.12-slim: it downloads
geo_score.pyfrom a pinned release, checks it againstSHA256SUMS, scores once, keeps
geo-score.jsonas an artifact and handsgeo-score.xmlto GitLab as a JUnit report, so a check at
zero shows as a failed test.examples/ci/README.mdgives the one directive that Bitbucket Pipelines,
Azure Pipelines, Jenkins and CircleCI each need to show the--junitfile. - Where to report what. A Benchmark correction form disputes a leaderboard row's score, gate
reason, label or not-scored reason, with a fresh--json-outreport and the tool version; a Tool
problem form covers the CLI, the Action, the MCP server, pip/pipx/uvx, the MCPB bundle, the Claude
Code plugin, the Gemini CLI extension and Docker. The issue chooser links the troubleshooting guide and
the MCP troubleshooting section, and a newSUPPORT.mdsays which form takes what. The forms applied
benchmark,engineandrubric, labels the repository did not have, so those issues arrived
unlabelled:.github/labels.jsonnow lists every label, andscripts/sync_labels.pycarries the list
to GitHub (a dry run by default;--applycreates and updates,--checkexits 1 when a label is
missing, and it never deletes one). It prints a proxy's host and port, never the credentials in
HTTPS_PROXY. - Benchmark: a re-runnable retest study.
benchmark/retest.pytakes two or more run directories
(results.json plus raw/) and records each run's tool, date, Python, declared vantage and pinning. For
each pair it reports the shares identical, within ±2 and within ±5, ICC(A,1) and SEM, band changes and
gate-cap flips, and per check the tier flips and the unobservable rate in each run. It keeps one record
per site per run, so every figure can be recomputed;--checkverifies the committed files
(validate.pyrule 32).benchmark/run.pygains--out DIR(a run for the study that leaves the
published files alone),--urls-from-dir DIR(each site re-reads an earlier run's pages, through the
CLI's--urls-fromreader) and--vantage LABEL, recorded in results.json asmeasured_frombeside
pythonandpinned_from. The operator declares the label; nothing looks an address up. - Benchmark: the Chinese-market gap, taken apart.
benchmark/stats.jsongainschina_gap_detail,
drawn after the existing bootstraps from the same seed, so no published interval moves: the gap by
audience language beside the sector-label gap, with their cross-tab (37/27/13/240); four matched sector
families, with no figure below 10 sites; the raw-point gap split into rule-read and heuristic checks
(68% heuristic, an upper bound on what a heuristic misreading could explain);p3.knowledge-graph
unobservable on 33 of 50 Chinese-language sites against 76 of 267 others; and the gap with that check
imputed as 0 (20 points) or as 4 (19 points). - Benchmark: flat downloads.
benchmark/results.csvhas one row per listed site (its fields,
sample_n,sample_source, the not-scored reason and the error, then one column per check holding its
points, blank where the check was not observable);benchmark/checks.csvhas one row per scored site
and check.run.pywrites both,--readme-onlyincluded; the leaderboard's Dataset markup and
llms.txt link them;validate.pyrule 35 checks every cell against results.json and raw/. There is no
robots.csv yet: raw/ does not record the status each user-agent got. - Benchmark: what each row rests on. Every row's real sample size (
sample_n,sample_source) is
published, read from raw/ for the committed rows. The 69 rows scored on fewer than 4 pages are marked
thin sample in the benchmark README and on both pages, which can filter on them, and stats.json puts
the median without them (62) beside the headline 56 (validate.pyrule 33). A table computed from
raw/ shows how common each practice is:llms.txtserved on 137 of 317 sites, complete Organization
and WebSite markup on 98, BreadcrumbList on a nested page on 82 of 251, a sitemap withlastmodon
161, a personal byline on at least half the pages on 67. The Web Almanac 2025llms.txtfigure sits
beside it, with a line saying the populations differ. - Benchmark: how it fetches, and how to correct a row. "How the benchmark fetches" states the
per-site request budget, computed from the CLI's own user-agent lists and constants (at most 67
requests to the site and 20 elsewhere, counted before redirects and retries), and says that crawler
user-agents are sent from the operator's address, not the vendors'. "Corrections" links the correction
form: a confirmed tool error ships as a patch release, and published scores change only when a site is
measured again.run.pyfiles a CLI error that starts with
bot challenge served to a browser user-agentunder a not-scored reason of its own, and the Not
scored list is grouped by reason. A row measured with 1.5.0 or later keeps the CLI's error kind
(error_kindin results.json) and is filed by it:http_*as refused,timeout,slowanddeadline
as timed out, anddns,refused,tls,droppedandno_responseas connection failed. Rows
measured with 1.3.0 are still filed by their message. When a site and itswwwhost both fail, the row
keeps the error of the attempt that got further (it kept thewwwhost's), and an IP address is not
retried onwww. - Weight sensitivity (
benchmark/sensitivity.py,SENSITIVITY.md+ zh). It first recomputes all 317
scores from raw/ (0 mismatches), then re-scores them under equal weights, an illustrative
evidence-basis weighting, each check left out, and 300 seeded draws. The rank order holds (Spearman
ρ ≥ 0.988 under every weighting); the bands move more: equal weights change 39 sites' bands, the
evidence-basis weighting 24, the draws a median of 26. Per-check tier spreads and item-rest
correlations follow, and it ends with input for v1.2. It adds withmath.fsum, so every Python writes
the same files. No v1.1 weight changes.validate.pyrule 36. - Where the band cuts fall (
benchmark/percentiles.py,percentiles.json, a table in the benchmark
README,rubric/bands-analysis.md+ zh). Of the 317 sites, 85%, 58%, 34% and 10% reach Early,
Growing, Solid and Leading. The analysis sets the v1.1 cuts against three criteria (achievable as Core
Web Vitals defines it: Leading is reached by 10%, with no margin; 168 of 317 sites sit within the
retest SEM's noise of a cut; 48 of Early's 88 sites are gate-capped) and lists candidates for v1.2. The
v1.1 cuts are unchanged.validate.pyrule 37. - The gate on every Python.
scripts/check.sh --allruns the three CLI suites and the MCP
conformance check on every Python from 3.8 to 3.14, and on 3.15 when it is installed;--strictfails
withuv python install Xwhen a needed Python is missing instead of skipping it. Both combine with
--live, and an unknown argument exits 2. The gate runs everyscripts/test_*.pyand
benchmark/test_*.pyonpython3, so a new suite joins it by its name, and under--liveit runs
scripts/sync_labels.py --checkwhenghis installed and logged in. - New tests.
scripts/test_ci_examples.pyruns the copyable CI examples with bash against fixture
sites.scripts/test_repo_meta.pychecks the forms against the labels, the README badge against its
recorded run, and the links in the chooser, the forms,SUPPORT.mdand the PR template.
benchmark/test_benchmark.pytests retest.py, the qualifiers, the canonical URLs, the China-gap
suppression, the sample sizes, the fetch budget, the CSVs, the sensitivity and percentile arithmetic
and the validity figures, offline.scripts/test_packaging.pynow checks that every tracked file
naming a release is onerelease.py --bumprewrites (the changelogs, benchmark/raw/, dist/,
docs/*.html and .github/validate.py are exempt), that every release-wheel URL's tag and file name
equal__version__, that validate.yml installs the ruff and jsonschema versions check.sh pins and
pins its actions by commit, that the Dockerfile's base is pinned by digest, and that the pyproject
classifiers cover 3.8 to 3.14 while validate.yml and RELEASING.md name the same two ends.
cli/test_geo_score.pyadds seven inputs that crashed 1.4.0's parsers and a seeded property test that
sends 200 random pages and robots.txt files through the parsers andrun(), and it runs the MCP
address guard's connect-time pinning, its refusal of a name that rebinds to a private address and its
CONNECT branch for HTTPS through a proxy, code no test had run before.cli/test_geo_score_e2e.py
injects faults (dropped connections, 429, fixed statuses, answers that trickle in), counts every
request the good fixture makes, and holdsdiffto golden files.
Changed
- Docs: what an API run cannot see, and the OpenRouter example. The README and the watch methodology (a
new §7) list the surfaces these APIs cannot reach and say that a citation is presence, not merit. The
OpenRouter example engine is nowdeepseek-via-openrouter, and the docs say an OpenRouter engine measures
the model answering over OpenRouter's search, not the vendor's own;countryis applied by OpenAI, Claude and
Perplexity and ignored by Gemini and OpenRouter. - The headline names the gate. Each gate at zero is named with the first sentence of its evidence, in
the default view and in--brief:
GATE AT ZERO — g.ssr: 0 of 8 sampled pages carry substantive body text …; the score cannot exceed 40 while a gate is at zero.
It used to say onlyGATE CAPPED — a gate check scored zero. - One order of gaps in every view. The default view,
--briefand--sharerank gaps the same way:
gates at zero first, then the points the next tier adds, then rubric order. Each row prints the number
that belongs to the condition beside it, for example
+4 next tier (+9 to full) Self-contained answer passages needs: …. Before, it printed the full
headroom next to the next tier's condition.--sharenow names the gate: "robots.txt disallows AI
retrieval crawlers", "AI retrieval crawlers are refused" or "crawlers get no server-rendered body
text". Before, it said a crawler "cannot reach the content at all", whichever gate it was. A content
gap now reads(+N at the next tier). The MCPscore_sitesummary follows too: it opens with the same
headline when a gate is at zero and lists up to six gaps in the same order, each with what the next
tier adds. It used to list the six checks with the lowest share of their points, bonus checks among
them. - The report footer.
Needs judgement:is split in two.Needs off-site search or judgement:lists
p3.listings, p3.mentions and p4.question-coverage.Not observable this run:lists the rest, each with
a short reason, for examplep4.cn-engines (not a Chinese-language site). The report and--brief
print a legend for ✓ ◐ ✗ ⊘. The rubric and troubleshooting links point at this release's tag
(blob/v1.5.0/…), as each check'sdoc_urlalready did. The terminal layout is not a contract. - One fetch wave after the sample. Once the sampled pages are in, every request whose URL is known
goes out at once and the checks read the answers from the run's memo: the sitemaps,llms.txt,
llms-full.txt,ai.txtand/.well-known/ai.txt(and the same files under a given path), the
Organization logo, up to 8sameAsprofiles, the Chinese crawlers and the 30 retrieval probes. The
Wikidata and Wikipedia lookups run beside them, one brand candidate at a time as before, so they get
the same lookups. Requests to the site still share 8 connections; other hosts get their own. On the
test fixtures, with 0.2 s added to every answer, a run went from 18–29 rounds of requests one after
another to 11–12. The wave itself moves no score and asks for the same URLs as 1.4.0, except
/.well-known/ai.txt, which is now requested even when/ai.txtexists. The retries it adds for the
site's files, the logo and the Chinese crawlers do move scores, and are under Fixed. --helpstands alone. It names the command you ran (geo_score.pyorgeo-score), listsdiff,
watchandmcpin the usage line, and explains the URL argument and-q. It ends with five
examples, the exit codes (watch's too) and links to the rubric and the troubleshooting guide for this
release. Withoutgeo_watch.py,mcp,watchand--asksay which of them needs it, and show the
release wheel'suvxandpipxcommands before the clone option.- The SVG badge of a gate-capped score reads
40/100 Early (gate capped), as the endpoint JSON does. diffis now a subcommand, likewatchandmcp. To score a host literally nameddiff, write
https://diff.- The scorer's documents. guide/troubleshooting.md (EN and zh) has a section for each error kind,
which the CLI's messages link to at the release tag, a "Why did my score move" section, and "Each page
is read in its own language" in place of the section on Chinese characters and English words.
rubric/v1.1.md (EN and zh) says underg.robotsandg.reachablewhich ten crawlers are read and how,
and underg.ssrthat a bot-challenge page is not a sampled page and never reads as tier 0; these are
clarifications, and no tier or points change. reference/ai-crawlers.md (verified 2026-09-28) says that
RETRIEVAL_UASis tested against its retrieval table, that the probes are unverified vendor strings
sent from the auditor's address, that the Claude-SearchBot and Claude-User strings are built from
Anthropic's documented product token, and that Amazonbot stays a retrieval crawler until Amazon's new
description of it is reviewed. rubric/open-questions.md (EN and zh) opens two questions for v1.2: #10,
whether a robots.txt that answers 5xx on every try is a full disallow, and #11, whether a sample with
no page carrying three figures should leavep2.sourced-statsout rather than score 3. - Install from the release wheel. The uvx, pipx and MCP install lines download
https://github.com/jianruntech/geo-score/releases/download/vX.Y.Z/geo_score-X.Y.Z-py3-none-any.whl,
so they need neither git nor a build step: the README install bullet, the Cursor, VS Code, VS Code
Insiders and LM Studio buttons, the Claude Code line, the standardmcpServersconfig, every
per-client block in guide/mcp.md and the MCPB watch-init hint. The git install
(git+https://github.com/jianruntech/geo-score@vX.Y.Z) stays on a line labelled "From source (needs
git)".release.py --bumprewrites the version in the wheel's file name as well as the tag. - The gate and the release. By default the gate looks for Python 3.8, 3.12 and 3.14 (it used to look
for 3.8, 3.12 and 3.13, never 3.14) and skips one uv cannot find, with its install hint;--strict
requires 3.8 and 3.14. It pins ruff 0.15.15 and jsonschema 4.23.0 (the newest jsonschema that installs
on Python 3.8): the PATH ruff runs only when it is the pinned version, otherwiseuvx ruff@0.15.15,
from uv's cache before the network, and python3's own jsonschema is used when it is the pinned
version. Offline with a cold cache it falls back to the PATH ruff and python3's jsonschema, as before
the pin. Under--all, the Python after the newest supported one is reported as optional, and--all
without uv says that only python3 runs.
release.py --executestops at once unlessghis logged in, runscheck.sh --live --strict(it ran
--livebefore), and stops when the built wheel's file name differs from the one the install lines
download. - validate.yml sets
permissions: {contents: read}, pins actions/checkout v7.0.1 and
actions/setup-python v7.0.0 by commit SHA, setspersist-credentials: false, installs the pinned ruff
and jsonschema, tests 3.8 and 3.14, and runs onubuntu-24.04, the newest image setup-python still
builds 3.8 for. The Dockerfile's base image ispython:3.12-slimpinned by its image-index digest
(3.12.14-slim-trixie, 2026-09-28). pyproject.toml has classifiers for Python 3.8 to 3.14 and
"3 :: Only". - Copyable CI.
examples/ci/github-actions.ymlpinsactions/checkout,actions/upload-artifactand
github/codeql-action/upload-sarifto commit SHAs with the release as a comment, sets
persist-credentials: false, passes step outputs throughenv:, and says how to hold the Action
still between runs (a release tag, or that release's commit).examples/ci/watch-weekly.ymlpins
actions/checkoutandactions/setup-pythonthe same way and keeps its credentials, because it
pushes the run file. - The README badge is a recorded run. It is regenerated from one run on the leaderboard page, kept
indocs/aiv-badge.json(2026-09-28, geo-score 1.4.0): 93/100 Leading, and its alt text names that
date and version. The old badge read 87 and was committed by hand. The scorer did not change. - Contributing. The pull request template asks for the last lines of
bash scripts/check.sh, both
changelogs, English and Chinese kept in step, the benchmark regeneration and disclosure of an
employer's site, and uses only absolute links. It and CONTRIBUTING.md name every file a new rubric
version needs: the next version's rubric in both languages and as JSON, both evidence tables and a
public audit scored with it. CONTRIBUTING.md also lists the two new test suites. - The leaderboard has one address. Its canonical, og:url, og:image, hreflang, llms.txt and sitemap
URLs point at the maintained mirror, https://www.jianruntech.com/leaderboard and /leaderboard/zh,
instead of the GitHub Pages copy frozen at 229 sites;GEO_SITE_BASEstill overrides.
scripts/check_live.pydefaults to the mirror and reads the site count, median and date from the page
itself.validate.pyrule 30. - The CLI-vs-hand-audit headline compares like with like. On the checks both scored, the CLI reads
0.4 points lower on average (per site from 6 lower to 9 higher, mean absolute difference 5.2;
Bland-Altman bias −0.4, 95% limits −12.8 to +12.0, n=5), kept inbenchmark/validity.jsonas
common_checksbeside avalidation_setblock. The full-score figure, 3.6 points lower, comes second:
the hand audits also scorep3.listings,p3.mentionsandp4.question-coverage, which the CLI
leaves out. VALIDITY.md (EN + zh) adds three caveats: the five sites are the v1.0→v1.1 calibration
sites, so this is development data; they score 69–87, with one Chinese site and none capped; and the
audit files do not record the rater. Both READMEs lead with the like-for-like figure, and
validate.pyrule 23 holds it andcommon_checksto the per-site rows. The 1.3.0 entry stays as
released. - The Chinese-market gap is decomposed rather than quoted whole, in benchmark/README.md, on the
leaderboard pages and in both READMEs: the 20-point gap states its definition (the five China
sectors), the language gap (20 points, 95% CI 13–25), the range within matched sector families (6 to
30 points) and the heuristic upper bound (68%).validate.pyrules 31 and 38. - REPRODUCIBILITY.md says what the ±5 / 96% was measured with: 229 sites on 2026-09-09, with a 1.1.x
build before the 1.1.2 variance fixes. It adds ICC(A,1) 0.974, SEM 3.09 and 18 band changes, and its
tables are generated by retest.py.reproducibility.jsonmoves to format 2 and keeps those runs under
legacy. The pages' footer and both READMEs name the date and the build. - Evidence notes, metadata only (no weight or tier changes):
p2.question-intentrecords that
through 1.4.x the CLI's Chinese heading lexicon lacked the task and explanation forms its English
lexicon matches, andp3.knowledge-graphthat the CLI queries Wikidata and Wikipedia but not Baidu
Baike. The rubric README and the evidence table link the sensitivity study, and the calibration record
the bands analysis.
Fixed
-
The sample is documents. A file that a sitemap or a link lists (Markdown such as
agents.md, feeds,
JSON, PDFs, media) is never a sampled page, and listing pages (pagination, category, tag, author, series
and archive pages) are drawn only after every content page. Direction: either way, toward what the site's
content reads. On the 1.3.0 leaderboard, 5 sites had a file in their sample and 39 had two or more
listing pages; baseus.com had been scored onagents.mdalone. -
p1.llms-txt: no answer is not observed. An/llms.txtthat gets no answer after its retries now
leaves the denominator instead of scoring 0. Direction: 0 to not observed (38 sites on the first 1.5.0
pass read 0 from a request that never got an answer). -
A handshake cut off in the middle is a dropped connection. An SSL EOF during the handshake is now kind
droppedand is retried, nottls, which is not. Atlserror on a Python whose TLS library cannot speak
TLS 1.3 says so: macOS's/usr/bin/python3is built on LibreSSL 2.8, and many sites require TLS 1.3. -
The leaderboard tries network failures again.
benchmark/run.pytries a site whose failure was a network
fault (a dropped connection, a timeout, a TLS error, a 429 or a 5xx) once more at the end of the run. The
1.5.0 leaderboard is run on Python 3.12: a first pass on the system Python lost 68 sites to TLS errors. -
g.robotsandg.reachablescore the 10 retrieval crawlers ofreference/ai-crawlers.md, and only
those. Claude-User, Googlebot and Amazonbot are added. GPTBot, ClaudeBot and a made-up
Google-Extended probe are removed. Each crawler is sent with its vendor-documented user-agent string.
g.robotsnow chooses each crawler's rules as RFC 9309 2.2.1 says: a crawler's own group replaces*.
A leading byte-order mark no longer hides the first group. Scores move in three ways:- Up for sites that block only training crawlers or opt-out tokens. On the leaderboard, measured
2026-09-28: together.ai, lonelyplanet.com, loom.com, stackoverflow.blog, webmd.com, airbnb.com and
khanacademy.org leaveg.robots0. airbnb.com stays capped byg.reachable, because it returns 403
to every crawler user-agent. - Up for a
*Disallow alongside named Allows for the retrieval crawlers. - Down where a site blocks Claude-User, Googlebot or Amazonbot, or where a byte-order mark hid a
blocking group. healthline.com stays at 0, now for Amazonbot.
Evidence now names the blocked crawlers and their group. Training crawlers and opt-out tokens
disallowed by name are listed as "not scored".g.reachableevidence lists the agents it did not
reach. - Up for sites that block only training crawlers or opt-out tokens. On the leaderboard, measured
-
A transient fault is retried before it moves a gate. robots.txt, the homepage and the sitemap are
fetched once per run, with retries.g.reachableprobes retry once on no answer, 429 or 5xx, and honour
Retry-After up to 5 s. A crawler that gets no answer on any page (a timeout or a dropped connection on
every try) is not counted: from one address that is mostly the network, or a firewall dropping a crawler
user-agent sent from an address its vendor does not use, which the real crawler never meets. The tier is
read over the crawlers that answered; tier 0 ("most retrieval user-agents are blocked") needs HTTP
refusals from more than half of all ten; with fewer than half answering, the check is not observed. A crawler answered
on some pages counts as reached. This rule came from 1.5.0's first leaderboard run, which read silence as
refusal and gate-capped 31 sites that a browser and most crawlers reached (pulumi.com went from 99 to 40);
on the five hand audits it takesg.reachable's exact agreement from 0.4 to 0.8. -
A homepage that cannot be fetched is not a missing entity. When the homepage is in the sample but
fails,p1.organization,p3.sameasandp3.knowledge-graphread not observed instead of 0 if the
other pages lack the entity: it usually sits on the homepage. Direction: 0 to not observed, on runs
where the homepage failed (nextjs.org's hand-audit comparison had read all three as 0 after one failed
fetch). A robots.txt that cannot be read after 3 tries is now not observed and leaves the
denominator, instead of reading as "No robots.txt". One that answers 429 or 5xx on every try keeps tier
3 and says so. The direction is either way: down where a dropped robots.txt fetch hid a Disallow, up
where a transient 429 or reset on a crawler probe read as a block. Five rows of the 1.3.0 leaderboard
took theirg.robotsfrom a failed fetch ("No robots.txt (HTTP 0)"): clickpaas.com, codecademy.com,
continue.dev, helium10.com and xcmg.com. -
The site's files, the Organization logo and the Chinese crawlers are retried. 1.4.0 asked once for
llms.txt,llms-full.txt,ai.txt, the logo and the home page as each Chinese crawler, so one
timeout, reset, 429 or 5xx read as a file that is not there, a logo that does not resolve or a crawler
that was not reached. The files and the logo are now retried like the sitemaps (up to 3 tries) and the
Chinese crawlers like the retrieval probes (one retry); a request still unanswered after its tries reads
as before, and a definite answer is asked once, so a healthy site gets no extra request. Direction: up,
where a transient fault used to read as absent, inp1.llms-txt,b.llms-full,b.ai-txt, the logo
part ofp1.organizationandp4.cn-engines. Two of seven live runs through a proxy had lost
p1.llms-txtto a single timeout. -
A bot-challenge page is not scored. When a bot-protection service (AWS WAF, Akamai, Cloudflare,
Fastly, PerimeterX, DataDome or Imperva) sends the CLI's browser user-agent a challenge page instead of
the page, that page is dropped from the sample like a failed fetch. A site that serves nothing but
challenges used to be scoredg.ssr0 and capped at 40, and the leaderboard published that as "content
needs JavaScript". It is now not scored: it exits 2 with
bot challenge served to a browser user-agent (<vendor>): a static fetch cannot measure this site, and
under--jsonit gives the error kindbot_challenge. Direction:g.ssrgoes from 0 to not scored
for such sites: live on 2026-09-30, ups.com (Akamai), which the 1.3.0 leaderboard shows at 16, capped,
got the bot-challenge error. A site that serves challenges on only some pages is scored on the real
pages left, and its notes count the challenges: booking.com (AWS WAF, 20 and capped on the 1.3.0
leaderboard) got the bot-challenge error on one run on 2026-09-30, and on two others that day was
scored on the pages left. openai.com, which the 1.3.0 leaderboard lists as "all requests returned
403", answers with a Cloudflare challenge and now gets the bot-challenge error. A JavaScript shell that
only loads a vendor's sensor script still scoresg.ssr0. HTTP error pages are now decompressed when
gzipped, so a gzipped challenge is recognised. -
A report scored on fewer pages says so.
notesgains
Scored on N of M sampled pages; K returned errors, J were bot challengeswhen pages were dropped.
When discovery found fewer pages than asked for, it gains
Sampled N of the 8 pages asked for: no more were found (sample source: …).schema/report.v2.json
(sampled_urls) has always asked for this. The terminal report prints the same lines. No score
changes. -
p2.sourced-statscounts a source only beside its figure. A figure is attributed only when its own
block of the content (a paragraph, list item, table cell, definition, caption or quotation), or the next
block, holds a link to a source off the site, a<cite>, a source phrase, or a footnote reference whose
note links to a source. The content is the first<main>or<article>, or else the page without its
header, navigation, footer and asides. Links to the site's own hosts, to its ownsameAsprofiles and
to social profiles or share buttons are not sources. 1.4.0 took any two outbound links,<cite>or
<blockquote>elements or source words anywhere in a page's main HTML as the source of all its figures,
so a social icon, a partner logo or a "Research" menu item did it. Fewer things are figures: a year, a
version number, a dotted date, a day of the month, a#reference, an identifier (seven or more digits
in a row, or a number thatICP备,ICP证,公安备,网安备or号introduces, or that号follows: filing,
account and phone numbers), and a price (a sum written the way only a price is, such as$29/mo,
68元/月,$15 per useror¥68起, one after a price word such as每月,售价,起价or低至, or one that
is all its block says, such as a pricing card's$29). Any other sum of money is still a figure, and so
is one with a magnitude after it ($4.5 billion,3000亿元). The tiers keep their page-fraction
structure: 7 needs three quarters of the pages that carry three or more figures to attribute most of
them with a link, 5 needs half with any attribution, 3 any attributed figure; a sample with no such page
keeps 3. The evidence gives the counts. The direction is down. Re-scored with 1.5.0 on 2026-09-30 on the
pages their 1.3.0 leaderboard reports sampled, sentry.io and seeyon.com read 3 (7 on the 1.3.0
leaderboard) and mingdao.com 0 (7); mistral.ai, whose pages are French, is not measured (5; see page
language below). On the five hand-audited sites, re-scored on their audited pages the same day, the CLI
now matches the auditor's tier on 4 of 5, against 1 of 5 in VALIDITY.md (measured with 1.3.0, when it
read higher on the other four); on mingdao.com it reads 0 against the auditor's 3. -
Each page is read in its own language. A sampled page's language comes from its
<html lang>, or
else itsContent-Languageheader, unless the visible text contradicts it. With neither, it comes from
the script (kana is Japanese, hangul Korean, Han alone Chinese) and, for Latin-script text, from a few
common words (English, German, French, Spanish, Portuguese). Code samples (<pre>,<code>) are not
read for this, and a page declared Chinese, Japanese or Korean keeps its declaration while it carries
30 or more such characters. Through 1.4.x one Chinese page switched every page of a run to the Chinese
thresholds, a Japanese site was labelledzh-CN, and German or French headings were read with the
English lexicon. The checks that move, and which way:p2.answer-passagesgoes up. Each page is counted in its own unit: characters (50–200) on Chinese,
Japanese and Korean pages, words (25–120) everywhere else. English pages in a sample whose homepage
is Chinese are now counted in words (airwallex.com's had scored 0 of 8), and kana and hangul now
count as characters. The evidence says which unit applied to how many pages.g.ssrgoes up on Japanese and Korean pages, whose kana and hangul now count toward the 120 floor (a
Japanese fixture went from 0, which capped its score at 40, to 5). It also goes up on English pages
whose main text carries more than 20 Han characters, such as a language menu. The "content needs
JavaScript" reading moves the same way.p2.question-intent: pages in a language without a heading lexicon (the CLI has English and Chinese)
leave the count. A site with no page left is not measured ("no question-intent lexicon for ja")
instead of being read with the English lexicon. It goes up for Chinese task, explanation and
comparison headings (创建你的第一个应用,工作流的工作原理,价格与版本对比), with noun compounds
such as管理团队and处理器and the product specs对比度(contrast ratio) and原理图
(schematic) left out, and for English comparison headings (compare, vs, versus, compared).p2.sourced-statsandp2.named-authormove toward not measured. Source phrases and written "By …"
bylines are read only on English and Chinese pages. A link, a<cite>, a footnote, and an author in
structured data orrel="author"are read in any language. The check leaves the denominator instead
of scoring 0 in two cases: no page carrying figures could be read for a source phrase and the reading
is below 5, or no page could be read for a byline.p4.cn-enginesleaves the denominator on Japanese and Korean sites. It is observed when a sampled
page is Chinese or the host name ends in.cn, no longer on any page heavy with Han characters. It
also leaves the denominator on a site with no Chinese page whose address merely contains.cn, such
aswww.cnn.comorwww.cnbc.com: 1.4.0 tested for.cnanywhere in the address. No leaderboard
site is one of these.p4.answer-shapegoes down on some pages because the paragraph-length limit is now set per page.
English pages in a sample with a Chinese homepage are held to 90 words again, and Japanese paragraphs
count their kana.p3.knowledge-graph: Wikidata and Wikipedia are searched in English and in the most common page
language other than English, when Wikipedia has a large edition in it. Japanese sites are no longer
searched in Chinese. A brand name is still searched in two languages at most.
-
audience_languageis the most common page language, as a BCP 47 primary tag (en,ja,de, …).
Chinese stayszh-CN, andundmeans no page's language could be told. A sample with a Chinese
homepage and English pages is nowen. The field was already a BCP 47 string. The description in
schema/report.v2.json now says what it holds. -
An unreachable site fails fast and says why. A typo or a dead host took from 20 seconds to 4
minutes and ended in raw urllib text. The CLI now resolves the host once before a run (not when a proxy
carries the request, since the proxy resolves it), and a name that does not exist fails at once with
cannot resolve <host>: check the spelling, or your DNS/proxy settings (HTTPS_PROXY). A name that does
not resolve, a refused connection and a failed TLS handshake are not retried. The final error is one
sentence that names the kind of failure, the proxy when one carried the request (host and port, never
credentials) and the troubleshooting section for it, pinned to the release tag. Input that is not an
http or https URL with a host (whitespace, another scheme, a port outside 1–65535) is a usage error,
exit 2; localhost, IP literals and dotless intranet names are still scored. A site that answers is
scored exactly as before. -
A slow server can no longer hold a run.
--timeoutlimits each read, not their sum: a server that
sent one byte every few seconds held a fetch, and a CI job or an MCP worker with it, for hours. One
request, redirects included, now takes at most 30 s or twice--timeout, whichever is longer, and then
counts as kindslow, without a retry. One scoring run takes at most 300 s: after that no request
starts, one in flight stops as kinddeadline, and the checks that needed it are not observed, a gate
included, rather than scored 0. The HTML scanners that retried every unclosed tag to the end of the
page (80 KB of<script>xtook 1.8 s, and a 4 MB page would have taken hours) read a page in one
pass, with the same results. A site that answers within these limits scores as before. -
Malformed markup can no longer crash a run.
{"@type": {"@id": …}}, a link, Sitemap line or logo
URL such ashttp://[::1, a lone surrogate escape in a JSON-LD name, or JSON-LD nested 2000 deep ended
a run with a traceback and exit 1, which CI reads as "below--fail-under". Only a string names a type
now, a malformed link reads as no link, a lone surrogate becomes?(an emoji still comes through
whole), and JSON-LD is flattened without recursion. Any other unexpected error prints
geo-score internal error (<Type>) while scoring <url>; please report it at …/issuesand exits 2;
GEO_SCORE_DEBUG=1shows the traceback. No score changes. -
--samplebelow 1 is now a usage error (exit 2). Before, 0 crashed and a negative value silently
scored fewer pages. The Action'ssampleinput takes a whole number of 1 or more. -
cli/README.md said the GitHub Action regenerates the badge "if you point
json-outat a path". It
never did. The section now describes only what exists:--badge,--badge-json, the Action'sbadge
input and the shields.io recipe. -
The distance to the next band reads
1 point to Solid, not1 points to Solid, in the report and in
--brief. -
Saving a paid run is all or nothing. An answer holding half a surrogate pair (half an emoji, which
a provider can send) stopped the save half way and left a truncated run file: the paid run was lost,
andwatch report,watch diffand the MCP tools failed on the latest run until the file was deleted
by hand. A run is now written in full to a hidden temporary file, synced to disk and only then linked
into place under a run id no other file has (<id>-2and on when the id is taken), so a failed save
leaves no partial file and the earlier runs as they were. Text UTF-8 cannot carry is stored as\u
escapes, which load back as the same text. -
Half a surrogate pair no longer crashes
watch run,watch report,watch askorwatch diff.
Text output shows it as a\uescape, on stdout and in an-ofile alike, and
watch report --format jsonandwatch ask --format jsonprint such answers. -
The engines' tuning keys are checked.
max_tokens,max_searchesandmax_resultsmust be
positive integers andtoola string;watch check, and every command that reads the config, now
refuses anything else, including a string such as"2048", which 1.4.0 accepted. A value pasted into
max_tokensused to reachint(), and the error saved in the run file quoted it; the new error never
shows the value.runs,max_callsand the tuning keys accept2.0, as JSON Schema does. -
MCP
score_sitesays why it refuses a host. A name that does not resolve gets
cannot resolve <host>: check the spelling, or your DNS/proxy settings, with a note, when a proxy
applies, thatscore_siteresolves every host on the machine running the server to check it is
public; such a host used to be called private. A host that resolves to a private, loopback or reserved
address getsurl must be on a public host: <host> resolves to a private, loopback or reserved address. -
The MCP rubric resource's
sourcelinks to the rubric at the release tag instead ofmain, and the
mcpentry ingeo_score.py watch --helpnames the installed command,geo-score-mcp. -
Installing needs no git. The documented
uvxandpipxroutes and the one-click MCP buttons no
longer fail on a machine that has uv but no git ("Git operation failed"), and no longer build from
source on first start. guide/mcp.md (EN + zh) no longer saysscore_sitetakes about 20 seconds, a
figure that was never measured, and neither do the leaderboard pages ("about twenty seconds"), the
README cover image or the demo image's description. How long a run takes depends on the site and the
network, and the report'selapsed_srecords it. -
The 'Anything else' CI snippet ran the CLI twice, so the report you kept could differ from the run
that set the exit code, and it did not verify the download. It now downloadsgeo_score.pyand
SHA256SUMSfrom a pinned release and scores once with--json-out,--junitand--fail-under.
The checksum and the run are chained with&&, so a file that does not match stops the step before
scoring even in a shell withoutset -e(Azure Pipelines script steps) orpipefail. The GitLab job
and the weekly watch workflow check each file against its ownSHA256SUMSline, and the weekly
compare step no longer runs after a failed checksum. -
The pull request template's Scope link was relative, which gives a 404 in a PR body. CONTRIBUTING.md
told rubric contributors to addrubric/v1.1.md, the version already released; it now names the next
one. -
The leaderboard's groups claim only what the 1.3.0 data supports. While results.json was measured
with a tool older than 1.5.0, the robots.txt group carries a qualifier: the 1.3.0 reading counted
GPTBot, ClaudeBot and Google-Extended as retrieval crawlers, so the 23 include sites that block only
training crawlers or an opt-out token. The JavaScript group carries one too: that reading could not
tell a bot-challenge page from a JavaScript shell. Both appear in benchmark/README.md, on both pages
and in llms.txt, and drop out by themselves once the leaderboard is re-run with 1.5.0 or later. The
403 group now reads "the server refuses crawler user-agents (unverified probes from the auditor's
address)", and the pages no longer say the China gap "is not content quality. It is llms.txt,
Organization schema…": its two largest per-check differences are heuristics. -
The benchmark README no longer claims "the same eight-page sample": it says up to eight pages per
site, and that 69 sites were scored on fewer than 4 pages (55 of them on a single page). -
Leaderboard Dataset markup.
readinesswas described as the "0-100 normalised score"; it is raw
points, andnormalisedis the percentage.variableMeasuredlists normalised, readiness,
observable_max, band and gate_capped, and the download, which pointed at a blob/main page, is
raw.githubusercontent.com pinned tov<version>, with the two CSVs beside it. The licence stays MIT
for now.validate.pyrule 34. -
None of the benchmark entries above re-scores a site: results.json is still the 1.3.0 run of
2026-09-27, and no check id or check reading changes with them.
Verify downloads: shasum -a 256 -c SHA256SUMS.
中文更新日志:https://github.com/jianruntech/geo-score/blob/v1.5.0/CHANGELOG.zh-CN.md