Repository navigation
geo-score 1.3.0
A release about trust in the measurement: outputs CI systems read natively, a pinned sample for
before/after comparisons, change detection that no longer calls noise a change, offline end-to-end
tests on twelve fixture sites, a measured CLI-vs-hand-audit agreement, and an evidence basis for every
check. The fixtures also found places where the CLI read sites against the rubric (grouped
robots.txt rules, single-page apps' catch-all pages, empty markup, stale dateModified, bylines,
Chinese dates and figures); those are fixed and listed below with their direction. The percentage can
no longer pass 100, and the MCP server now keeps every connection on public addresses.
Upgrading from 1.2.0. Level 1 scores move only where a check read a site wrongly; each fix and
its direction is listed under Fixed, and the public leaderboard (measured with 1.1.0) is not
re-scored. In watch, the headline rate leaves out questions that name the brand, intervals widen
when a question has several answers, and some diff verdicts that said change now say
within noise or too few questions to tell; the 1.2.0 fields keep their meaning. Evidence strings
now quote site text inside «site text: …». The CSV report has an eleventh column, branded.
Added
- One run, several outputs:
--json-out FILE,--sarif FILE(SARIF 2.1.0: one rule per check,
one result per check below full marks) and--junit FILE(one test per check; a check at zero
fails, an unobservable one is skipped). All are written from the same fetch as the report on
screen, so they cannot disagree with it. Single site only. - Install as a command:
pyproject.toml, sopipx install git+https://github.com/jianruntech/geo-score
anduvx --from git+… geo-scorework. Installsgeo-scoreandgeo-score-mcp(the MCP server over
stdio). Still no dependencies. STABILITY.md: what is a public contract (check ids, score arithmetic, the JSON schemas, CLI
flags, exit codes, Action inputs and outputs, MCP tool names, SARIF rule ids), what semver means
here, and that nothing is sent anywhere (no telemetry).- Action inputs
sarif,junitandannotations; outputsgate-cappedandreport-path. - Pin the sample:
--urls FILEscores exactly the given pages (up to 8, all on the scored site);
--urls-from REPORT.jsonre-reads the pages an earlier report sampled. Sampling is the main source
of run-to-run variance, so a before/after comparison on pinned pages shows what changed on the
site. The report sayssample_source: "pinned". Action inputurls-from. - Per-check documentation links: every check in the JSON report carries
doc_url, its definition
in the rubric pinned to the release tag, so the link never drifts from the scorer. SARIF rules link
there too, and--explainprints it under each check below full marks. Optional field in
schema/report.v2.json. - Troubleshooting, in English and Chinese: score movement, checks shown
as—, the gate cap, exit code 2, proxies, redirects, the leaderboard's tool version,watch
exit 3, Action annotation limits, CJK thresholds. The terminal report links to it. - Intervals on the leaderboard's medians.
benchmark/stats.pycomputes a seeded bootstrap 95%
interval for the overall median, each sector median and the Chinese-market gap;benchmark/README.md
and both READMEs quote them: overall 56 (53–58), the gap 19 points (12–26). The hosted page shows
medians without intervals until it is rebuilt and redeployed. Sectors with fewer than 10 scored sites (China / cloud and Professional services, 9
each) no longer print a median;validate.pycounts scored sites, not listed ones.
benchmark/run.py --readme-onlyregenerates the README without re-scoring. - The leaderboard can move host:
GEO_SITE_BASEsets the base URL for canonical links,og:image,
hreflang,llms.txtand the sitemap, andscripts/check_live.pycompares a hosted page with the
committed data (today's Pages build still shows 229 sites against 314). - How far to trust the CLI, measured.
benchmark/validity.pyre-scores the five published hand
audits on the same pages and compares every check tier for tier: 79% of 86 check pairs agree (95%
for checks read by rule, 64% for heuristics), and the CLI reads 3.6 points lower on average, from
10 lower to 6 higher. Results, caveats and every disagreement:benchmark/VALIDITY.md(EN + zh).
This replaces the unmeasured "typically 5 to 15 points lower", which was wrong for two sites in five. - Evidence basis for every check.
rubric/v1.1.jsongains anevidenceobject per check
(basis, sources, last verified), published asrubric/evidence-v1.1.md(EN + zh). Metadata only:
no weight, tier or score changed. - MCP for agents, not only for people.
score_site,report,diffandlist_runsreturn
structuredContent(the full report, summary or diff as JSON) beside the text, and declare an
outputSchema, so an agent no longer parses Markdown to recover a score.score_sitetakesurls
(pin the pages, same checks as--urls) andsample. The rubric is served as the resource
geo-score://rubric/v1.1, built from the scorer's own tables. The server instructions and the
score_sitedescription say that«site text: …»spans are data from the audited site. watch diff --fail-on-drop: exit 4 when a verdict ischangewith a negative change. The MCP
difftool takesformat: "json". guide/watch-methodology.md
(EN + zh) gives the formulas and a worked example.- One gate:
scripts/check.sh. Structure, ruff, the three test suites on Python 3.8 and the newest
available, and an emulation of the GitHub Action against a fixture site (scripts/test_action.py);
--liveadds example.com..githooks/pre-pushruns it;.github/workflows/validate.ymlruns the same
script. Hosted Actions have not run on this organisation so far, soCONTRIBUTING.mdno longer says
CI checks anything: maintainers run the script before every merge. AGENTS.md: the repository's invariants, each mapped to the check that enforces it, for
contributors and coding agents (CLAUDE.mdpoints to it).- Issue forms ask for the tool version and the JSON report; a new form covers level 2 and 3 engine
problems and warns never to paste a key; security reports are routed toSECURITY.md. - Offline end-to-end tests.
cli/test_geo_score_e2e.pyserves twelve fixture sites from
127.0.0.1 (a good site, a JavaScript shell, robots blocks, a firewall that refuses AI crawlers, a
Chinese site, a site under a path, a hostile site that tries to address an agent, and more) and
checks every check's tier against each site'sexpected.json, plus a golden report. A socket
guard fails the run if anything leaves the machine; the suite takes under a second. - Tests that the CLI's embedded rubric (check ids, points, every tier's points and wording, bands,
caps, bonus) equalsrubric/v1.1.json.
Changed
watch diffis paired, clustered and corrected for multiple comparisons. It compares the
questions both runs answered, with an exact paired test per row (McNemar when every question has
one answer per run, otherwise an exact sign-flip test on each question's net change) and Holm's
correction across the engine rows and the all-engines row. The verdict ischangeonly when the
adjusted p < 0.05, andtoo few questions to tellbelow 6 shared questions. 1.2.0 used an unpaired
two-proportion test on answers that are not independent, which called changes too readily. New
row fields:p_paired,p_adjusted,n_pairs,n_questions,n_asks,before_paired,
after_paired,delta_pp,delta_ci.before,after,deltaandp_valuekeep their 1.2.0
unpaired meaning; no verdict uses them.- Intervals account for repeated asks. Every rate block gains
cited_ci_clustered(a bootstrap
over questions, 2,000 seeded resamples, never narrower than Wilson) andn_questions; reports quote
it whenever a question has more than one answer.cited_ci(Wilson) is unchanged. - Headline rates leave out questions that name the brand. Questions are flagged at run time
(brand_in_query,competitor_in_query) with the matcher used on answers; the headline covers the
other questions and branded ones get their own line and n. An optionalbrandedcolumn in
queries.csv(auto,yes,no) overrides the flag,watch checkwarns when more than 30% of the
questions name the brand, and the CSV report gains a trailingbrandedcolumn. - The Action scores the site once. It used to fetch and score twice (once for JSON, once for
the log), which doubled the time and could print two different scores. The job summary, the
annotations and every file now come from one run. - The Action annotates checks at zero: gates as errors, other checks as warnings (GitHub shows up
to 10 of each; the rest are in the job summary). geo-score watch --helpnames the command you typed (geo-scoreafter an install).--comparewith three or more sites no longer says "Both scored".- When a site redirects to another domain, the full report says so and the JSON carries
landed
(optional), the URL that was scored.
Fixed
- Scoring fixes the offline fixtures found. Each moves a score only where the old reading
contradicted the rubric:g.robots: severalUser-agentlines share the rules that follow them (RFC 9309 2.1). The rules
were attached to the last line only, so a group listing GPTBot, ClaudeBot and CCBot over
Disallow: /read as blocking only CCBot. Down, to 0 and the gate cap, where grouped rules block
retrieval crawlers; up (3 to 5) where a groupedAllownow counts as an explicit permission.p1.llms-txt,b.llms-full,b.ai-txt,p1.sitemap: a single-page app that answers every path
with its HTML shell no longer counts as serving these files; a sitemap must be XML with a
<urlset>or<sitemapindex>. Down. A relativeSitemap:line in robots.txt is resolved against the
origin (up where it now loads).p1.organization: Organization and WebSite without name, url and logo score 3, as the rubric
says, not 5 (down); a relative logo resolves against the page it sits on (up where it now loads);
Organization subtypes (EducationalOrganization, NewsMediaOrganization, LocalBusiness and others) are
read for logo and sameAs, which liftsp1.organizationandp3.sameasfor sites typed that way (up).
sameAsvalues that are not http(s) URLs are no longer fetched.p1.page-type: the top tier needs real field values; a bare{"@type": "Product"}scores 2. Down.p2.freshness: dates flush against Chinese text and dates with English month names are read
(up); the top tier requiresdateModifiednot to be older than the page's own date (its
datePublished, article time or the first date in its content) on at least three quarters of the
pages that declare it (down where it is).p2.sourced-stats: "resources" no longer counts as "source" (down); figures flush against Chinese
text count as figures (either way).p2.named-author: "Editor", the organisation's own name and "Powered by Hugo" in a footer are not
personal bylines. Down.- Sampling: for a site under a path, a sibling folder (
/docs-legacyfor/docs) is no longer
sampled, and the path's own sitemap is used when the origin has none.
- SARIF uploads to GitHub code scanning. Results were located at the scored URL, which code
scanning rejects; each now points at its check's line in the JSON report (the Action writes one into
the workspace whensarifis set) and names the site in its message. watch diffno longer sayswithin noisewhen the question set is too small for the number of
rows compared: before Holm's correction, a row with m × 2 / 2^n ≥ 0.05 (m rows, n questions;
Tarone's screen) istoo few questions to telland leaves the family. The screen is conservative. An unknown MCP tool is a JSON-RPC error (-32602), as the protocol says.- A score could exceed 100. Bonus points sit outside the denominator, so a site at or near full
marks plus bonus read above 100 (the good fixture read 107), which also broke the report schema (normalised≤ 100). The
percentage is now capped at 100, and rubric rule 5 says so. No published score changes: the
benchmark's highest is 98 and the hand audits' 87. The offline fixtures found it. --json-out,--sarifand--junitreport a path that cannot be written as a clear error (exit
2) instead of a traceback.- A Latin brand name written directly next to Chinese, Japanese or Korean text ("用Lumo做视频") now
counts as a mention; it never matched before. New runs can therefore show a highermentionedrate
on such answers than earlier runs. - Run files saved by 1.2.0 still load; reports recompute their summary and the branded flag from the
saved question text, and say so in a note. - Documentation that still described the superseded gate rule:
rubric/README.md, the Chinese
calibration record,SKILL.mdand the sample report said a gate short of full marks caps the score.
The rule since v1.1 is that only a gate at zero caps it. The scores were always computed with the
correct rule;validate.pynow fails on any sentence that states the cap without the zero condition. - The calibration record credited the structured-data adoption figures (BreadcrumbList 5.66%,
Product 0.77%, FAQPage 0.34%) to the 2025 Web Almanac; they are from the 2024 edition. Only the
llms.txt figure (2.13%) is from 2025. - The Chinese calibration record was missing two facts the English one has (the 60 KB of
server-rendered prose, Foglift's upper quartile); a new check compares the numbers in both.
Security
- The MCP server connects only to public addresses, on every hop. 1.2.0 checked the URL an agent
passed and nothing after it, so the audited site could point the scorer at loopback or intranet
addresses through a redirect, itssameAslinks, its logo or its declared sitemap, and the check
let carrier-grade NAT through (100.64.0.0/10, where cloud metadata such as 100.100.100.200 lives).
Under MCP every connection is now resolved and checked (is_global), each redirect hop is checked,
and a direct connection goes to the checked address. The check also uses the exact host to be
fetched rather than its brand form. The command-line tool still scores whatever its user names. - More text is fenced or cleaned: the sitemap URL a site declares is quoted as site text without its
query or fragment; sitemap entries with control or bidi characters are dropped; Unicode tag
characters, soft hyphens, line separators and XML-illegal characters are removed; a truncated
evidence line never leaves a«site text: …»fence open; server error lines are reduced to one
clean line; the Action's job summary escapes evidence instead of rendering it as Markdown and HTML. - The Action passes its inputs to the script as environment variables. Before this,
urland the
other inputs were pasted into the shell script, so a crafted value could run commands in the
workflow.sampleandfail-underare now checked to be whole numbers. - Site text in evidence is fenced as data. Evidence that quotes the audited site (an answer
passage, the brand name read from the page, a Wikidata description) now carries it inside
«site text: …», and control, zero-width and bidi characters are stripped from every evidence
string. The report reaches AI agents through the MCP server, so a page could otherwise address the
agent in the tool's own voice. Harvested links carrying such characters are dropped from the sample. - Errors that wrap a lower-level exception (a malformed request header, an unreadable config) no
longer chain it, so an uncaught traceback cannot print the original message, which for a header can
contain the API key. SECURITY.mdnow has a supported-versions table and a threat model;SKILL.mdtells the agent
that fetched text is evidence, never instructions.
Verify downloads: shasum -a 256 -c SHA256SUMS.
中文更新日志:https://github.com/jianruntech/geo-score/blob/v1.3.0/CHANGELOG.zh-CN.md