Skip to content

geo-score 1.5.0

Choose a tag to compare

@vortwang vortwang released this 02 Oct 16:39
· 59 commits to main since this release

A release about reading sites the way the rubric says, and about failing CI for the right reason. The two
crawler gates now read only the ten retrieval crawlers, choose each crawler's robots.txt group as RFC 9309
does and retry a request that got no answer before they count it; a bot-challenge page is no longer scored
as a page that needs JavaScript; every page is read in its own language; and a figure counts as sourced
only when its source sits beside it. These fixes move scores, and each one below names its check and its
direction. For CI there are --fail-on-gate and --assert (a gate at zero caps a score at 40, and a site
that would otherwise score 40 or more reads exactly 40, which --fail-under 40 lets through),
geo-score diff and --baseline to compare two reports check by check, and a shields.io endpoint badge.
A site that cannot be scored fails fast and says why in one line, in JSON under --json, and a scoring
run can no longer hang. Around the CLI: a schema and typo warnings for the watch config, installs from the
release wheel without git, a gate that pins its tools, CI examples beyond GitHub, forms for disputing a
leaderboard row, and a benchmark that can be re-measured and taken apart. The leaderboard itself is still
the 1.3.0 run of 2026-09-27.

Upgrading from 1.4.0. Level-1 scores move only where 1.4.0 read a site wrong, and every such fix is
under Fixed with its check and direction: the gates (g.robots, g.reachable, and g.ssr for sites
behind a bot challenge, with Japanese or Korean pages, or with English pages that carry more than 20 Han
characters), the checks that read page text (p2.answer-passages, p2.question-intent,
p2.sourced-stats, p2.named-author, p4.answer-shape, p4.cn-engines and p3.knowledge-graph), and
the checks whose file, logo or crawler request is now retried (p1.llms-txt, b.llms-full, b.ai-txt,
p1.organization and p4.cn-engines). Three invocations exit differently: a site that serves only
bot-challenge pages exits 2 instead of being scored; an unexpected error exits 2 in one line instead of
a traceback that exited 1; and --sample below 1 exits 2, where 1.4.0 crashed on 0 and scored fewer
pages on a negative value. diff is now a subcommand, so a host named diff is written
https://diff. audience_language can hold languages other than en and zh-CN, and the report gains
only the optional elapsed_s.

Added

  • A glossary and a Chinese CLI reference. guide/glossary.md (and guide/glossary.zh-CN.md) defines the 49
    terms the README, the reports and the guides use, each with where it appears and which page goes deeper;
    cli/README.zh-CN.md is the CLI reference in Simplified Chinese, with a language switch at the top of both.
  • --fail-on-gate and --assert EXPR, in the CLI and as the Action inputs fail-on-gate and
    assert. A gate at zero caps the score at 40: a site that would otherwise score 40 or more reads
    exactly 40 and passes --fail-under 40 (fail-under: 40). --fail-on-gate fails any capped site: it
    exits 1 when any g.* check scores 0, and names the gate and its evidence; it is the same as
    --assert 'g.*>0'. --assert holds one check, or an fnmatch glob of checks, to a value, e.g.
    g.robots=max or p1.llms-txt>=4, and is repeatable. A check the run could not observe is skipped,
    never failed. A malformed expression, or a glob that names no check, exits 2.
    Both are opt-in and change no score, and their results go to stderr, never into the report.
  • geo-score diff BEFORE AFTER [--format text|md|json] [--fail-on-drop] compares two level-1 reports
    check by check. Each check shows its points before and after, its max, and a state: up, down,
    same, became unobservable, became observable, or unobservable when neither report observed it.
    The comparison also shows the score, band and gate cap before and after, names each gate that reached
    or left 0, and lists the pages added to or removed from the sample. When the two samples differ, it
    says so, reads the totals as ±5 (benchmark/REPRODUCIBILITY.md) and names g.reachable as the least
    stable check. A check that became unobservable is never shown as a drop to 0. --format json follows
    the new schema/diff.v1.json. Two reports of different rubric versions exit 2, and reports from
    different releases, or of two different targets, are compared with a warning on stderr. diff needs
    only geo_score.py.
  • --baseline REPORT on a scoring run scores the baseline's pages again (as --urls-from does) and
    prints the same comparison after the report. It cannot be combined with --urls or --urls-from.
    Under --json, stdout carries only the report and the comparison is not printed: write the report
    with --json-out and run geo-score diff on the two files to see it. --fail-on-drop still exits 4
    and names each drop on stderr.
  • --fail-on-drop (with --baseline, or on diff) exits 4 when a check lost points, a gate reached
    0 or the band fell, and prints one stderr line per reason. It is opt-in, as watch diff --fail-on-drop
    is, so no existing invocation changes its exit code. A failed --fail-under, --fail-on-gate or
    --assert still exits 1.
  • Action: a baseline input, delta and dropped-checks outputs, and a "Changes since baseline"
    section in the job summary, with the check ids read from the baseline escaped. The baseline may be the
    same file as json-out.
  • A badge that stays current. --badge-json FILE writes the badge as shields.io
    endpoint
    JSON:
    {"schemaVersion": 1, "label": "AIV readiness", "message": "71/100 Solid", "color": "#2f8f52", "cacheSeconds": 86400}.
    The SVG and the JSON take their message and colour from one helper, so the two cannot disagree. Host
    the file anywhere and https://img.shields.io/endpoint?url=<its URL> draws the badge. The Action gets
    a badge input: a path ending in .json gets the endpoint JSON, and any other path gets the SVG. A
    badge-path output says where the file went.
  • A JSON error instead of nothing. Under --json, every exit 2, usage errors included, prints
    {"error": {"kind", "message", "target"}, "tool": "geo-score-cli/X.Y.Z"} on stdout, where stdout used
    to be empty; diff --format json does the same. The new schema/error.v1.json names the kinds: dns,
    refused, tls, timeout, dropped and no_response when nothing answered; http_<code>,
    no_content, slow and bot_challenge when the site answered; deadline when the run ran out of
    time; not_public under the MCP server's address guard; and usage, rubric_mismatch (reports of two
    rubric versions given to diff or --baseline), error and internal. A minor release may add
    kinds; read an unknown one as error.
  • --timeout SECONDS (default 15, as before; up to 300): how long each request may wait for the
    server before it counts as unanswered.
  • elapsed_s in the JSON report (optional, schema/report.v2.json): how many seconds the run took
    on the machine that ran it. The footer prints it too (· took 23.4 s). It describes the run, not the
    site, and diff ignores it.
  • A schema for geo-score-watch.json, and a warning for every key nothing reads.
    schema/watch-config.v1.json (JSON Schema draft 2020-12) describes the file, and watch init now
    writes "$schema" as its first key, so an editor checks and completes it. A mistyped key used to
    vanish without a word: budget for budget_usd meant no budget, and max_call for max_calls the
    default cap of 120. watch check names every unknown key at every level (top, brand,
    competitors[i], engines[i], prices.<id>), with the nearest known key when one is close, as in
    unknown key 'budget' (did you mean 'budget_usd'?); it is ignored, and flags a prices entry that
    names no engine. watch run and the other commands that read the config print the same warnings on
    stderr; the MCP status tool lists them, and run shows them with its plan and after a saved run. A
    key named like a credential (api_key, OPENAI_KEY, api_token, client_secret, credentials …)
    is pointed at environment variables instead, and its value is never printed. Warnings only: no exit
    code changes.
  • CI beyond GitHub. examples/ci/gitlab-ci.yml is a GitLab job on python:3.12-slim: it downloads
    geo_score.py from a pinned release, checks it against SHA256SUMS, scores once, keeps
    geo-score.json as an artifact and hands geo-score.xml to GitLab as a JUnit report, so a check at
    zero shows as a failed test. examples/ci/README.md gives the one directive that Bitbucket Pipelines,
    Azure Pipelines, Jenkins and CircleCI each need to show the --junit file.
  • Where to report what. A Benchmark correction form disputes a leaderboard row's score, gate
    reason, label or not-scored reason, with a fresh --json-out report and the tool version; a Tool
    problem
    form covers the CLI, the Action, the MCP server, pip/pipx/uvx, the MCPB bundle, the Claude
    Code plugin, the Gemini CLI extension and Docker. The issue chooser links the troubleshooting guide and
    the MCP troubleshooting section, and a new SUPPORT.md says which form takes what. The forms applied
    benchmark, engine and rubric, labels the repository did not have, so those issues arrived
    unlabelled: .github/labels.json now lists every label, and scripts/sync_labels.py carries the list
    to GitHub (a dry run by default; --apply creates and updates, --check exits 1 when a label is
    missing, and it never deletes one). It prints a proxy's host and port, never the credentials in
    HTTPS_PROXY.
  • Benchmark: a re-runnable retest study. benchmark/retest.py takes two or more run directories
    (results.json plus raw/) and records each run's tool, date, Python, declared vantage and pinning. For
    each pair it reports the shares identical, within ±2 and within ±5, ICC(A,1) and SEM, band changes and
    gate-cap flips, and per check the tier flips and the unobservable rate in each run. It keeps one record
    per site per run, so every figure can be recomputed; --check verifies the committed files
    (validate.py rule 32). benchmark/run.py gains --out DIR (a run for the study that leaves the
    published files alone), --urls-from-dir DIR (each site re-reads an earlier run's pages, through the
    CLI's --urls-from reader) and --vantage LABEL, recorded in results.json as measured_from beside
    python and pinned_from. The operator declares the label; nothing looks an address up.
  • Benchmark: the Chinese-market gap, taken apart. benchmark/stats.json gains china_gap_detail,
    drawn after the existing bootstraps from the same seed, so no published interval moves: the gap by
    audience language beside the sector-label gap, with their cross-tab (37/27/13/240); four matched sector
    families, with no figure below 10 sites; the raw-point gap split into rule-read and heuristic checks
    (68% heuristic, an upper bound on what a heuristic misreading could explain); p3.knowledge-graph
    unobservable on 33 of 50 Chinese-language sites against 76 of 267 others; and the gap with that check
    imputed as 0 (20 points) or as 4 (19 points).
  • Benchmark: flat downloads. benchmark/results.csv has one row per listed site (its fields,
    sample_n, sample_source, the not-scored reason and the error, then one column per check holding its
    points, blank where the check was not observable); benchmark/checks.csv has one row per scored site
    and check. run.py writes both, --readme-only included; the leaderboard's Dataset markup and
    llms.txt link them; validate.py rule 35 checks every cell against results.json and raw/. There is no
    robots.csv yet: raw/ does not record the status each user-agent got.
  • Benchmark: what each row rests on. Every row's real sample size (sample_n, sample_source) is
    published, read from raw/ for the committed rows. The 69 rows scored on fewer than 4 pages are marked
    thin sample in the benchmark README and on both pages, which can filter on them, and stats.json puts
    the median without them (62) beside the headline 56 (validate.py rule 33). A table computed from
    raw/ shows how common each practice is: llms.txt served on 137 of 317 sites, complete Organization
    and WebSite markup on 98, BreadcrumbList on a nested page on 82 of 251, a sitemap with lastmod on
    161, a personal byline on at least half the pages on 67. The Web Almanac 2025 llms.txt figure sits
    beside it, with a line saying the populations differ.
  • Benchmark: how it fetches, and how to correct a row. "How the benchmark fetches" states the
    per-site request budget, computed from the CLI's own user-agent lists and constants (at most 67
    requests to the site and 20 elsewhere, counted before redirects and retries), and says that crawler
    user-agents are sent from the operator's address, not the vendors'. "Corrections" links the correction
    form: a confirmed tool error ships as a patch release, and published scores change only when a site is
    measured again. run.py files a CLI error that starts with
    bot challenge served to a browser user-agent under a not-scored reason of its own, and the Not
    scored list is grouped by reason. A row measured with 1.5.0 or later keeps the CLI's error kind
    (error_kind in results.json) and is filed by it: http_* as refused, timeout, slow and deadline
    as timed out, and dns, refused, tls, dropped and no_response as connection failed. Rows
    measured with 1.3.0 are still filed by their message. When a site and its www host both fail, the row
    keeps the error of the attempt that got further (it kept the www host's), and an IP address is not
    retried on www.
  • Weight sensitivity (benchmark/sensitivity.py, SENSITIVITY.md + zh). It first recomputes all 317
    scores from raw/ (0 mismatches), then re-scores them under equal weights, an illustrative
    evidence-basis weighting, each check left out, and 300 seeded draws. The rank order holds (Spearman
    ρ ≥ 0.988 under every weighting); the bands move more: equal weights change 39 sites' bands, the
    evidence-basis weighting 24, the draws a median of 26. Per-check tier spreads and item-rest
    correlations follow, and it ends with input for v1.2. It adds with math.fsum, so every Python writes
    the same files. No v1.1 weight changes. validate.py rule 36.
  • Where the band cuts fall (benchmark/percentiles.py, percentiles.json, a table in the benchmark
    README, rubric/bands-analysis.md + zh). Of the 317 sites, 85%, 58%, 34% and 10% reach Early,
    Growing, Solid and Leading. The analysis sets the v1.1 cuts against three criteria (achievable as Core
    Web Vitals defines it: Leading is reached by 10%, with no margin; 168 of 317 sites sit within the
    retest SEM's noise of a cut; 48 of Early's 88 sites are gate-capped) and lists candidates for v1.2. The
    v1.1 cuts are unchanged. validate.py rule 37.
  • The gate on every Python. scripts/check.sh --all runs the three CLI suites and the MCP
    conformance check on every Python from 3.8 to 3.14, and on 3.15 when it is installed; --strict fails
    with uv python install X when a needed Python is missing instead of skipping it. Both combine with
    --live, and an unknown argument exits 2. The gate runs every scripts/test_*.py and
    benchmark/test_*.py on python3, so a new suite joins it by its name, and under --live it runs
    scripts/sync_labels.py --check when gh is installed and logged in.
  • New tests. scripts/test_ci_examples.py runs the copyable CI examples with bash against fixture
    sites. scripts/test_repo_meta.py checks the forms against the labels, the README badge against its
    recorded run, and the links in the chooser, the forms, SUPPORT.md and the PR template.
    benchmark/test_benchmark.py tests retest.py, the qualifiers, the canonical URLs, the China-gap
    suppression, the sample sizes, the fetch budget, the CSVs, the sensitivity and percentile arithmetic
    and the validity figures, offline. scripts/test_packaging.py now checks that every tracked file
    naming a release is one release.py --bump rewrites (the changelogs, benchmark/raw/, dist/,
    docs/*.html and .github/validate.py are exempt), that every release-wheel URL's tag and file name
    equal __version__, that validate.yml installs the ruff and jsonschema versions check.sh pins and
    pins its actions by commit, that the Dockerfile's base is pinned by digest, and that the pyproject
    classifiers cover 3.8 to 3.14 while validate.yml and RELEASING.md name the same two ends.
    cli/test_geo_score.py adds seven inputs that crashed 1.4.0's parsers and a seeded property test that
    sends 200 random pages and robots.txt files through the parsers and run(), and it runs the MCP
    address guard's connect-time pinning, its refusal of a name that rebinds to a private address and its
    CONNECT branch for HTTPS through a proxy, code no test had run before. cli/test_geo_score_e2e.py
    injects faults (dropped connections, 429, fixed statuses, answers that trickle in), counts every
    request the good fixture makes, and holds diff to golden files.

Changed

  • Docs: what an API run cannot see, and the OpenRouter example. The README and the watch methodology (a
    new §7) list the surfaces these APIs cannot reach and say that a citation is presence, not merit. The
    OpenRouter example engine is now deepseek-via-openrouter, and the docs say an OpenRouter engine measures
    the model answering over OpenRouter's search, not the vendor's own; country is applied by OpenAI, Claude and
    Perplexity and ignored by Gemini and OpenRouter.
  • The headline names the gate. Each gate at zero is named with the first sentence of its evidence, in
    the default view and in --brief:
    GATE AT ZERO — g.ssr: 0 of 8 sampled pages carry substantive body text …; the score cannot exceed 40 while a gate is at zero.
    It used to say only GATE CAPPED — a gate check scored zero.
  • One order of gaps in every view. The default view, --brief and --share rank gaps the same way:
    gates at zero first, then the points the next tier adds, then rubric order. Each row prints the number
    that belongs to the condition beside it, for example
    +4 next tier (+9 to full) Self-contained answer passages needs: …. Before, it printed the full
    headroom next to the next tier's condition. --share now names the gate: "robots.txt disallows AI
    retrieval crawlers", "AI retrieval crawlers are refused" or "crawlers get no server-rendered body
    text". Before, it said a crawler "cannot reach the content at all", whichever gate it was. A content
    gap now reads (+N at the next tier). The MCP score_site summary follows too: it opens with the same
    headline when a gate is at zero and lists up to six gaps in the same order, each with what the next
    tier adds. It used to list the six checks with the lowest share of their points, bonus checks among
    them.
  • The report footer. Needs judgement: is split in two. Needs off-site search or judgement: lists
    p3.listings, p3.mentions and p4.question-coverage. Not observable this run: lists the rest, each with
    a short reason, for example p4.cn-engines (not a Chinese-language site). The report and --brief
    print a legend for ✓ ◐ ✗ ⊘. The rubric and troubleshooting links point at this release's tag
    (blob/v1.5.0/…), as each check's doc_url already did. The terminal layout is not a contract.
  • One fetch wave after the sample. Once the sampled pages are in, every request whose URL is known
    goes out at once and the checks read the answers from the run's memo: the sitemaps, llms.txt,
    llms-full.txt, ai.txt and /.well-known/ai.txt (and the same files under a given path), the
    Organization logo, up to 8 sameAs profiles, the Chinese crawlers and the 30 retrieval probes. The
    Wikidata and Wikipedia lookups run beside them, one brand candidate at a time as before, so they get
    the same lookups. Requests to the site still share 8 connections; other hosts get their own. On the
    test fixtures, with 0.2 s added to every answer, a run went from 18–29 rounds of requests one after
    another to 11–12. The wave itself moves no score and asks for the same URLs as 1.4.0, except
    /.well-known/ai.txt, which is now requested even when /ai.txt exists. The retries it adds for the
    site's files, the logo and the Chinese crawlers do move scores, and are under Fixed.
  • --help stands alone. It names the command you ran (geo_score.py or geo-score), lists diff,
    watch and mcp in the usage line, and explains the URL argument and -q. It ends with five
    examples, the exit codes (watch's too) and links to the rubric and the troubleshooting guide for this
    release. Without geo_watch.py, mcp, watch and --ask say which of them needs it, and show the
    release wheel's uvx and pipx commands before the clone option.
  • The SVG badge of a gate-capped score reads 40/100 Early (gate capped), as the endpoint JSON does.
  • diff is now a subcommand, like watch and mcp. To score a host literally named diff, write
    https://diff.
  • The scorer's documents. guide/troubleshooting.md (EN and zh) has a section for each error kind,
    which the CLI's messages link to at the release tag, a "Why did my score move" section, and "Each page
    is read in its own language" in place of the section on Chinese characters and English words.
    rubric/v1.1.md (EN and zh) says under g.robots and g.reachable which ten crawlers are read and how,
    and under g.ssr that a bot-challenge page is not a sampled page and never reads as tier 0; these are
    clarifications, and no tier or points change. reference/ai-crawlers.md (verified 2026-09-28) says that
    RETRIEVAL_UAS is tested against its retrieval table, that the probes are unverified vendor strings
    sent from the auditor's address, that the Claude-SearchBot and Claude-User strings are built from
    Anthropic's documented product token, and that Amazonbot stays a retrieval crawler until Amazon's new
    description of it is reviewed. rubric/open-questions.md (EN and zh) opens two questions for v1.2: #10,
    whether a robots.txt that answers 5xx on every try is a full disallow, and #11, whether a sample with
    no page carrying three figures should leave p2.sourced-stats out rather than score 3.
  • Install from the release wheel. The uvx, pipx and MCP install lines download
    https://github.com/jianruntech/geo-score/releases/download/vX.Y.Z/geo_score-X.Y.Z-py3-none-any.whl,
    so they need neither git nor a build step: the README install bullet, the Cursor, VS Code, VS Code
    Insiders and LM Studio buttons, the Claude Code line, the standard mcpServers config, every
    per-client block in guide/mcp.md and the MCPB watch-init hint. The git install
    (git+https://github.com/jianruntech/geo-score@vX.Y.Z) stays on a line labelled "From source (needs
    git)". release.py --bump rewrites the version in the wheel's file name as well as the tag.
  • The gate and the release. By default the gate looks for Python 3.8, 3.12 and 3.14 (it used to look
    for 3.8, 3.12 and 3.13, never 3.14) and skips one uv cannot find, with its install hint; --strict
    requires 3.8 and 3.14. It pins ruff 0.15.15 and jsonschema 4.23.0 (the newest jsonschema that installs
    on Python 3.8): the PATH ruff runs only when it is the pinned version, otherwise uvx ruff@0.15.15,
    from uv's cache before the network, and python3's own jsonschema is used when it is the pinned
    version. Offline with a cold cache it falls back to the PATH ruff and python3's jsonschema, as before
    the pin. Under --all, the Python after the newest supported one is reported as optional, and --all
    without uv says that only python3 runs.
    release.py --execute stops at once unless gh is logged in, runs check.sh --live --strict (it ran
    --live before), and stops when the built wheel's file name differs from the one the install lines
    download.
  • validate.yml sets permissions: {contents: read}, pins actions/checkout v7.0.1 and
    actions/setup-python v7.0.0 by commit SHA, sets persist-credentials: false, installs the pinned ruff
    and jsonschema, tests 3.8 and 3.14, and runs on ubuntu-24.04, the newest image setup-python still
    builds 3.8 for. The Dockerfile's base image is python:3.12-slim pinned by its image-index digest
    (3.12.14-slim-trixie, 2026-09-28). pyproject.toml has classifiers for Python 3.8 to 3.14 and
    "3 :: Only".
  • Copyable CI. examples/ci/github-actions.yml pins actions/checkout, actions/upload-artifact and
    github/codeql-action/upload-sarif to commit SHAs with the release as a comment, sets
    persist-credentials: false, passes step outputs through env:, and says how to hold the Action
    still between runs (a release tag, or that release's commit). examples/ci/watch-weekly.yml pins
    actions/checkout and actions/setup-python the same way and keeps its credentials, because it
    pushes the run file.
  • The README badge is a recorded run. It is regenerated from one run on the leaderboard page, kept
    in docs/aiv-badge.json (2026-09-28, geo-score 1.4.0): 93/100 Leading, and its alt text names that
    date and version. The old badge read 87 and was committed by hand. The scorer did not change.
  • Contributing. The pull request template asks for the last lines of bash scripts/check.sh, both
    changelogs, English and Chinese kept in step, the benchmark regeneration and disclosure of an
    employer's site, and uses only absolute links. It and CONTRIBUTING.md name every file a new rubric
    version needs: the next version's rubric in both languages and as JSON, both evidence tables and a
    public audit scored with it. CONTRIBUTING.md also lists the two new test suites.
  • The leaderboard has one address. Its canonical, og:url, og:image, hreflang, llms.txt and sitemap
    URLs point at the maintained mirror, https://www.jianruntech.com/leaderboard and /leaderboard/zh,
    instead of the GitHub Pages copy frozen at 229 sites; GEO_SITE_BASE still overrides.
    scripts/check_live.py defaults to the mirror and reads the site count, median and date from the page
    itself. validate.py rule 30.
  • The CLI-vs-hand-audit headline compares like with like. On the checks both scored, the CLI reads
    0.4 points lower on average (per site from 6 lower to 9 higher, mean absolute difference 5.2;
    Bland-Altman bias −0.4, 95% limits −12.8 to +12.0, n=5), kept in benchmark/validity.json as
    common_checks beside a validation_set block. The full-score figure, 3.6 points lower, comes second:
    the hand audits also score p3.listings, p3.mentions and p4.question-coverage, which the CLI
    leaves out. VALIDITY.md (EN + zh) adds three caveats: the five sites are the v1.0→v1.1 calibration
    sites, so this is development data; they score 69–87, with one Chinese site and none capped; and the
    audit files do not record the rater. Both READMEs lead with the like-for-like figure, and
    validate.py rule 23 holds it and common_checks to the per-site rows. The 1.3.0 entry stays as
    released.
  • The Chinese-market gap is decomposed rather than quoted whole, in benchmark/README.md, on the
    leaderboard pages and in both READMEs: the 20-point gap states its definition (the five China
    sectors), the language gap (20 points, 95% CI 13–25), the range within matched sector families (6 to
    30 points) and the heuristic upper bound (68%). validate.py rules 31 and 38.
  • REPRODUCIBILITY.md says what the ±5 / 96% was measured with: 229 sites on 2026-09-09, with a 1.1.x
    build before the 1.1.2 variance fixes. It adds ICC(A,1) 0.974, SEM 3.09 and 18 band changes, and its
    tables are generated by retest.py. reproducibility.json moves to format 2 and keeps those runs under
    legacy. The pages' footer and both READMEs name the date and the build.
  • Evidence notes, metadata only (no weight or tier changes): p2.question-intent records that
    through 1.4.x the CLI's Chinese heading lexicon lacked the task and explanation forms its English
    lexicon matches, and p3.knowledge-graph that the CLI queries Wikidata and Wikipedia but not Baidu
    Baike. The rubric README and the evidence table link the sensitivity study, and the calibration record
    the bands analysis.

Fixed

  • The sample is documents. A file that a sitemap or a link lists (Markdown such as agents.md, feeds,
    JSON, PDFs, media) is never a sampled page, and listing pages (pagination, category, tag, author, series
    and archive pages) are drawn only after every content page. Direction: either way, toward what the site's
    content reads. On the 1.3.0 leaderboard, 5 sites had a file in their sample and 39 had two or more
    listing pages; baseus.com had been scored on agents.md alone.

  • p1.llms-txt: no answer is not observed. An /llms.txt that gets no answer after its retries now
    leaves the denominator instead of scoring 0. Direction: 0 to not observed (38 sites on the first 1.5.0
    pass read 0 from a request that never got an answer).

  • A handshake cut off in the middle is a dropped connection. An SSL EOF during the handshake is now kind
    dropped and is retried, not tls, which is not. A tls error on a Python whose TLS library cannot speak
    TLS 1.3 says so: macOS's /usr/bin/python3 is built on LibreSSL 2.8, and many sites require TLS 1.3.

  • The leaderboard tries network failures again. benchmark/run.py tries a site whose failure was a network
    fault (a dropped connection, a timeout, a TLS error, a 429 or a 5xx) once more at the end of the run. The
    1.5.0 leaderboard is run on Python 3.12: a first pass on the system Python lost 68 sites to TLS errors.

  • g.robots and g.reachable score the 10 retrieval crawlers of reference/ai-crawlers.md, and only
    those.
    Claude-User, Googlebot and Amazonbot are added. GPTBot, ClaudeBot and a made-up
    Google-Extended probe are removed. Each crawler is sent with its vendor-documented user-agent string.
    g.robots now chooses each crawler's rules as RFC 9309 2.2.1 says: a crawler's own group replaces *.
    A leading byte-order mark no longer hides the first group. Scores move in three ways:

    • Up for sites that block only training crawlers or opt-out tokens. On the leaderboard, measured
      2026-09-28: together.ai, lonelyplanet.com, loom.com, stackoverflow.blog, webmd.com, airbnb.com and
      khanacademy.org leave g.robots 0. airbnb.com stays capped by g.reachable, because it returns 403
      to every crawler user-agent.
    • Up for a * Disallow alongside named Allows for the retrieval crawlers.
    • Down where a site blocks Claude-User, Googlebot or Amazonbot, or where a byte-order mark hid a
      blocking group. healthline.com stays at 0, now for Amazonbot.

    Evidence now names the blocked crawlers and their group. Training crawlers and opt-out tokens
    disallowed by name are listed as "not scored". g.reachable evidence lists the agents it did not
    reach.

  • A transient fault is retried before it moves a gate. robots.txt, the homepage and the sitemap are
    fetched once per run, with retries. g.reachable probes retry once on no answer, 429 or 5xx, and honour
    Retry-After up to 5 s. A crawler that gets no answer on any page (a timeout or a dropped connection on
    every try) is not counted: from one address that is mostly the network, or a firewall dropping a crawler
    user-agent sent from an address its vendor does not use, which the real crawler never meets. The tier is
    read over the crawlers that answered; tier 0 ("most retrieval user-agents are blocked") needs HTTP
    refusals from more than half of all ten; with fewer than half answering, the check is not observed. A crawler answered
    on some pages counts as reached. This rule came from 1.5.0's first leaderboard run, which read silence as
    refusal and gate-capped 31 sites that a browser and most crawlers reached (pulumi.com went from 99 to 40);
    on the five hand audits it takes g.reachable's exact agreement from 0.4 to 0.8.

  • A homepage that cannot be fetched is not a missing entity. When the homepage is in the sample but
    fails, p1.organization, p3.sameas and p3.knowledge-graph read not observed instead of 0 if the
    other pages lack the entity: it usually sits on the homepage. Direction: 0 to not observed, on runs
    where the homepage failed (nextjs.org's hand-audit comparison had read all three as 0 after one failed
    fetch). A robots.txt that cannot be read after 3 tries is now not observed and leaves the
    denominator, instead of reading as "No robots.txt". One that answers 429 or 5xx on every try keeps tier
    3 and says so. The direction is either way: down where a dropped robots.txt fetch hid a Disallow, up
    where a transient 429 or reset on a crawler probe read as a block. Five rows of the 1.3.0 leaderboard
    took their g.robots from a failed fetch ("No robots.txt (HTTP 0)"): clickpaas.com, codecademy.com,
    continue.dev, helium10.com and xcmg.com.

  • The site's files, the Organization logo and the Chinese crawlers are retried. 1.4.0 asked once for
    llms.txt, llms-full.txt, ai.txt, the logo and the home page as each Chinese crawler, so one
    timeout, reset, 429 or 5xx read as a file that is not there, a logo that does not resolve or a crawler
    that was not reached. The files and the logo are now retried like the sitemaps (up to 3 tries) and the
    Chinese crawlers like the retrieval probes (one retry); a request still unanswered after its tries reads
    as before, and a definite answer is asked once, so a healthy site gets no extra request. Direction: up,
    where a transient fault used to read as absent, in p1.llms-txt, b.llms-full, b.ai-txt, the logo
    part of p1.organization and p4.cn-engines. Two of seven live runs through a proxy had lost
    p1.llms-txt to a single timeout.

  • A bot-challenge page is not scored. When a bot-protection service (AWS WAF, Akamai, Cloudflare,
    Fastly, PerimeterX, DataDome or Imperva) sends the CLI's browser user-agent a challenge page instead of
    the page, that page is dropped from the sample like a failed fetch. A site that serves nothing but
    challenges used to be scored g.ssr 0 and capped at 40, and the leaderboard published that as "content
    needs JavaScript". It is now not scored: it exits 2 with
    bot challenge served to a browser user-agent (<vendor>): a static fetch cannot measure this site, and
    under --json it gives the error kind bot_challenge. Direction: g.ssr goes from 0 to not scored
    for such sites: live on 2026-09-30, ups.com (Akamai), which the 1.3.0 leaderboard shows at 16, capped,
    got the bot-challenge error. A site that serves challenges on only some pages is scored on the real
    pages left, and its notes count the challenges: booking.com (AWS WAF, 20 and capped on the 1.3.0
    leaderboard) got the bot-challenge error on one run on 2026-09-30, and on two others that day was
    scored on the pages left. openai.com, which the 1.3.0 leaderboard lists as "all requests returned
    403", answers with a Cloudflare challenge and now gets the bot-challenge error. A JavaScript shell that
    only loads a vendor's sensor script still scores g.ssr 0. HTTP error pages are now decompressed when
    gzipped, so a gzipped challenge is recognised.

  • A report scored on fewer pages says so. notes gains
    Scored on N of M sampled pages; K returned errors, J were bot challenges when pages were dropped.
    When discovery found fewer pages than asked for, it gains
    Sampled N of the 8 pages asked for: no more were found (sample source: …). schema/report.v2.json
    (sampled_urls) has always asked for this. The terminal report prints the same lines. No score
    changes.

  • p2.sourced-stats counts a source only beside its figure. A figure is attributed only when its own
    block of the content (a paragraph, list item, table cell, definition, caption or quotation), or the next
    block, holds a link to a source off the site, a <cite>, a source phrase, or a footnote reference whose
    note links to a source. The content is the first <main> or <article>, or else the page without its
    header, navigation, footer and asides. Links to the site's own hosts, to its own sameAs profiles and
    to social profiles or share buttons are not sources. 1.4.0 took any two outbound links, <cite> or
    <blockquote> elements or source words anywhere in a page's main HTML as the source of all its figures,
    so a social icon, a partner logo or a "Research" menu item did it. Fewer things are figures: a year, a
    version number, a dotted date, a day of the month, a # reference, an identifier (seven or more digits
    in a row, or a number that ICP备, ICP证, 公安备, 网安备 or 号 introduces, or that 号 follows: filing,
    account and phone numbers), and a price (a sum written the way only a price is, such as $29/mo,
    68元/月, $15 per user or ¥68起, one after a price word such as 每月, 售价, 起价 or 低至, or one that
    is all its block says, such as a pricing card's $29). Any other sum of money is still a figure, and so
    is one with a magnitude after it ($4.5 billion, 3000亿元). The tiers keep their page-fraction
    structure: 7 needs three quarters of the pages that carry three or more figures to attribute most of
    them with a link, 5 needs half with any attribution, 3 any attributed figure; a sample with no such page
    keeps 3. The evidence gives the counts. The direction is down. Re-scored with 1.5.0 on 2026-09-30 on the
    pages their 1.3.0 leaderboard reports sampled, sentry.io and seeyon.com read 3 (7 on the 1.3.0
    leaderboard) and mingdao.com 0 (7); mistral.ai, whose pages are French, is not measured (5; see page
    language below). On the five hand-audited sites, re-scored on their audited pages the same day, the CLI
    now matches the auditor's tier on 4 of 5, against 1 of 5 in VALIDITY.md (measured with 1.3.0, when it
    read higher on the other four); on mingdao.com it reads 0 against the auditor's 3.

  • Each page is read in its own language. A sampled page's language comes from its <html lang>, or
    else its Content-Language header, unless the visible text contradicts it. With neither, it comes from
    the script (kana is Japanese, hangul Korean, Han alone Chinese) and, for Latin-script text, from a few
    common words (English, German, French, Spanish, Portuguese). Code samples (<pre>, <code>) are not
    read for this, and a page declared Chinese, Japanese or Korean keeps its declaration while it carries
    30 or more such characters. Through 1.4.x one Chinese page switched every page of a run to the Chinese
    thresholds, a Japanese site was labelled zh-CN, and German or French headings were read with the
    English lexicon. The checks that move, and which way:

    • p2.answer-passages goes up. Each page is counted in its own unit: characters (50–200) on Chinese,
      Japanese and Korean pages, words (25–120) everywhere else. English pages in a sample whose homepage
      is Chinese are now counted in words (airwallex.com's had scored 0 of 8), and kana and hangul now
      count as characters. The evidence says which unit applied to how many pages.
    • g.ssr goes up on Japanese and Korean pages, whose kana and hangul now count toward the 120 floor (a
      Japanese fixture went from 0, which capped its score at 40, to 5). It also goes up on English pages
      whose main text carries more than 20 Han characters, such as a language menu. The "content needs
      JavaScript" reading moves the same way.
    • p2.question-intent: pages in a language without a heading lexicon (the CLI has English and Chinese)
      leave the count. A site with no page left is not measured ("no question-intent lexicon for ja")
      instead of being read with the English lexicon. It goes up for Chinese task, explanation and
      comparison headings (创建你的第一个应用, 工作流的工作原理, 价格与版本对比), with noun compounds
      such as 管理团队 and 处理器 and the product specs 对比度 (contrast ratio) and 原理图
      (schematic) left out, and for English comparison headings (compare, vs, versus, compared).
    • p2.sourced-stats and p2.named-author move toward not measured. Source phrases and written "By …"
      bylines are read only on English and Chinese pages. A link, a <cite>, a footnote, and an author in
      structured data or rel="author" are read in any language. The check leaves the denominator instead
      of scoring 0 in two cases: no page carrying figures could be read for a source phrase and the reading
      is below 5, or no page could be read for a byline.
    • p4.cn-engines leaves the denominator on Japanese and Korean sites. It is observed when a sampled
      page is Chinese or the host name ends in .cn, no longer on any page heavy with Han characters. It
      also leaves the denominator on a site with no Chinese page whose address merely contains .cn, such
      as www.cnn.com or www.cnbc.com: 1.4.0 tested for .cn anywhere in the address. No leaderboard
      site is one of these.
    • p4.answer-shape goes down on some pages because the paragraph-length limit is now set per page.
      English pages in a sample with a Chinese homepage are held to 90 words again, and Japanese paragraphs
      count their kana.
    • p3.knowledge-graph: Wikidata and Wikipedia are searched in English and in the most common page
      language other than English, when Wikipedia has a large edition in it. Japanese sites are no longer
      searched in Chinese. A brand name is still searched in two languages at most.
  • audience_language is the most common page language, as a BCP 47 primary tag (en, ja, de, …).
    Chinese stays zh-CN, and und means no page's language could be told. A sample with a Chinese
    homepage and English pages is now en. The field was already a BCP 47 string. The description in
    schema/report.v2.json now says what it holds.

  • An unreachable site fails fast and says why. A typo or a dead host took from 20 seconds to 4
    minutes and ended in raw urllib text. The CLI now resolves the host once before a run (not when a proxy
    carries the request, since the proxy resolves it), and a name that does not exist fails at once with
    cannot resolve <host>: check the spelling, or your DNS/proxy settings (HTTPS_PROXY). A name that does
    not resolve, a refused connection and a failed TLS handshake are not retried. The final error is one
    sentence that names the kind of failure, the proxy when one carried the request (host and port, never
    credentials) and the troubleshooting section for it, pinned to the release tag. Input that is not an
    http or https URL with a host (whitespace, another scheme, a port outside 1–65535) is a usage error,
    exit 2; localhost, IP literals and dotless intranet names are still scored. A site that answers is
    scored exactly as before.

  • A slow server can no longer hold a run. --timeout limits each read, not their sum: a server that
    sent one byte every few seconds held a fetch, and a CI job or an MCP worker with it, for hours. One
    request, redirects included, now takes at most 30 s or twice --timeout, whichever is longer, and then
    counts as kind slow, without a retry. One scoring run takes at most 300 s: after that no request
    starts, one in flight stops as kind deadline, and the checks that needed it are not observed, a gate
    included, rather than scored 0. The HTML scanners that retried every unclosed tag to the end of the
    page (80 KB of <script>x took 1.8 s, and a 4 MB page would have taken hours) read a page in one
    pass, with the same results. A site that answers within these limits scores as before.

  • Malformed markup can no longer crash a run. {"@type": {"@id": …}}, a link, Sitemap line or logo
    URL such as http://[::1, a lone surrogate escape in a JSON-LD name, or JSON-LD nested 2000 deep ended
    a run with a traceback and exit 1, which CI reads as "below --fail-under". Only a string names a type
    now, a malformed link reads as no link, a lone surrogate becomes ? (an emoji still comes through
    whole), and JSON-LD is flattened without recursion. Any other unexpected error prints
    geo-score internal error (<Type>) while scoring <url>; please report it at …/issues and exits 2;
    GEO_SCORE_DEBUG=1 shows the traceback. No score changes.

  • --sample below 1 is now a usage error (exit 2). Before, 0 crashed and a negative value silently
    scored fewer pages. The Action's sample input takes a whole number of 1 or more.

  • cli/README.md said the GitHub Action regenerates the badge "if you point json-out at a path". It
    never did. The section now describes only what exists: --badge, --badge-json, the Action's badge
    input and the shields.io recipe.

  • The distance to the next band reads 1 point to Solid, not 1 points to Solid, in the report and in
    --brief.

  • Saving a paid run is all or nothing. An answer holding half a surrogate pair (half an emoji, which
    a provider can send) stopped the save half way and left a truncated run file: the paid run was lost,
    and watch report, watch diff and the MCP tools failed on the latest run until the file was deleted
    by hand. A run is now written in full to a hidden temporary file, synced to disk and only then linked
    into place under a run id no other file has (<id>-2 and on when the id is taken), so a failed save
    leaves no partial file and the earlier runs as they were. Text UTF-8 cannot carry is stored as \u
    escapes, which load back as the same text.

  • Half a surrogate pair no longer crashes watch run, watch report, watch ask or watch diff.
    Text output shows it as a \u escape, on stdout and in an -o file alike, and
    watch report --format json and watch ask --format json print such answers.

  • The engines' tuning keys are checked. max_tokens, max_searches and max_results must be
    positive integers and tool a string; watch check, and every command that reads the config, now
    refuses anything else, including a string such as "2048", which 1.4.0 accepted. A value pasted into
    max_tokens used to reach int(), and the error saved in the run file quoted it; the new error never
    shows the value. runs, max_calls and the tuning keys accept 2.0, as JSON Schema does.

  • MCP score_site says why it refuses a host. A name that does not resolve gets
    cannot resolve <host>: check the spelling, or your DNS/proxy settings, with a note, when a proxy
    applies, that score_site resolves every host on the machine running the server to check it is
    public; such a host used to be called private. A host that resolves to a private, loopback or reserved
    address gets url must be on a public host: <host> resolves to a private, loopback or reserved address.

  • The MCP rubric resource's source links to the rubric at the release tag instead of main, and the
    mcp entry in geo_score.py watch --help names the installed command, geo-score-mcp.

  • Installing needs no git. The documented uvx and pipx routes and the one-click MCP buttons no
    longer fail on a machine that has uv but no git ("Git operation failed"), and no longer build from
    source on first start. guide/mcp.md (EN + zh) no longer says score_site takes about 20 seconds, a
    figure that was never measured, and neither do the leaderboard pages ("about twenty seconds"), the
    README cover image or the demo image's description. How long a run takes depends on the site and the
    network, and the report's elapsed_s records it.

  • The 'Anything else' CI snippet ran the CLI twice, so the report you kept could differ from the run
    that set the exit code, and it did not verify the download. It now downloads geo_score.py and
    SHA256SUMS from a pinned release and scores once with --json-out, --junit and --fail-under.
    The checksum and the run are chained with &&, so a file that does not match stops the step before
    scoring even in a shell without set -e (Azure Pipelines script steps) or pipefail. The GitLab job
    and the weekly watch workflow check each file against its own SHA256SUMS line, and the weekly
    compare step no longer runs after a failed checksum.

  • The pull request template's Scope link was relative, which gives a 404 in a PR body. CONTRIBUTING.md
    told rubric contributors to add rubric/v1.1.md, the version already released; it now names the next
    one.

  • The leaderboard's groups claim only what the 1.3.0 data supports. While results.json was measured
    with a tool older than 1.5.0, the robots.txt group carries a qualifier: the 1.3.0 reading counted
    GPTBot, ClaudeBot and Google-Extended as retrieval crawlers, so the 23 include sites that block only
    training crawlers or an opt-out token. The JavaScript group carries one too: that reading could not
    tell a bot-challenge page from a JavaScript shell. Both appear in benchmark/README.md, on both pages
    and in llms.txt, and drop out by themselves once the leaderboard is re-run with 1.5.0 or later. The
    403 group now reads "the server refuses crawler user-agents (unverified probes from the auditor's
    address)", and the pages no longer say the China gap "is not content quality. It is llms.txt,
    Organization schema…": its two largest per-check differences are heuristics.

  • The benchmark README no longer claims "the same eight-page sample": it says up to eight pages per
    site, and that 69 sites were scored on fewer than 4 pages (55 of them on a single page).

  • Leaderboard Dataset markup. readiness was described as the "0-100 normalised score"; it is raw
    points, and normalised is the percentage. variableMeasured lists normalised, readiness,
    observable_max, band and gate_capped, and the download, which pointed at a blob/main page, is
    raw.githubusercontent.com pinned to v<version>, with the two CSVs beside it. The licence stays MIT
    for now. validate.py rule 34.

  • None of the benchmark entries above re-scores a site: results.json is still the 1.3.0 run of
    2026-09-27, and no check id or check reading changes with them.


Verify downloads: shasum -a 256 -c SHA256SUMS.
中文更新日志:https://github.com/jianruntech/geo-score/blob/v1.5.0/CHANGELOG.zh-CN.md