Repository navigation
Releases: jianruntech/geo-score
Release list
geo-score 1.8.0
Added
- How much of the answer your pages support, beside the cited rate. Where a provider documents where each
citation sits, every answer now records it (spans: start, end, URLs and kind, character offsets into the
answer as saved, at most 500), read only from documented fields: thestart_index/end_indexof OpenAI's and
OpenRouter'surl_citationannotations (where the citation sits), Perplexity's[n]marks when its search
results carry theidits documentation maps them to, Gemini'surl_citationannotations (Interactions API)
and grounding supports (generate_content), both measured in bytes as Google's reference says, and the text
of Claude's text blocks that carry citations. From them each answer getssupported_share: the share of the
answer a URL the brand owns supports, by characters, by sentences, and by characters weighted by their
sentence's position, e^(−position / sentences), after the position-adjusted word count of arXiv 2311.09735,
§2.2.1, Eq. 3 (characters for words; a sentence several sources support counts in full). A mark supports the
sentence it closes; a segment, its own characters. The report's new "Answer share your pages support" table
gives, per engine and next to its cited rate, never merged into it, the answers measured, the median of each
share and how many are above zero;summary.engines[].supported_sharecarries the same. An answer with no
citation, or with one that came without a usable position (an OpenAI-compatiblebase_urlthat strips
annotations looks the same), is not measured, never 0, and so is every Perplexity answer whose results carry
no id. It is notmention_share: that counts sentences naming the brand, this counts text its pages support. - Sources listed and sources used. Perplexity's adapter kept only the sources the text marks with
[n]and
dropped the rest; Gemini'sgenerate_contentlists grounding chunks and ties supports to some of them. Such
answers now recordsources_listedandsources_used, and a new table, "Sources listed and used in the
answer", gives per engine the sources listed and used per answer, used over listed, and the answers that list
and use a URL of yours (summary.engines[].sources).citations,cited, the rank and the headline rate read
exactly as before for the same responses, which a test holds against 1.7.1's reading. OpenAI, Claude,
OpenRouter and Gemini's Interactions API return one list, the citations tied to the text, so they are not
measured. - A run records what it measured in
recorded. A run saved before has none, and its report says the answer share
and the sources were not recorded rather than that the engines gave none; its summary is computed again when it
is reported, and the file is not rewritten. schema/watch.v1.json registers every new field, closed and capped;
guide/watch-methodology.md section 9 gives the definitions, which engines give positions and the papers behind
them; the glossary gains Answer share and Sources listed, sources used. Measurement only: rubric v1.1, every
tier and every score are unchanged. - Repeat consistency per question. When a run asks a question more than once on an engine, the report adds
"Repeat consistency of cited domains": the Jaccard overlap of the domains cited by every two answers that cite,
averaged per question, the median per engine, and the least consistent questions (summary.engines[].consistency).
Answers that cite nothing are left out; one ask per question reads not measured. Descriptive only:watch diff
does not read it. For scale, the methodology cites arXiv 2604.07585 §5 tables 4 and 6. - A precision hint in the plan.
watch run --dry-run(and the MCP run tool's dry run) prints, beside the
resolution, the worst-case standard error of one question's cited rate on one engine at the configuredruns
(0.5/√n), and that more asks add no questions. The methodology adds the external reference point of arXiv
2604.07585 §7 (at least 7 asks per prompt per day for brands, 8 for sources; one market, one study). - Interfaces are never pooled. The methodology (section 7) records that one product's web app, mobile app and
API cite only partly overlapping sources (arXiv 2607.15771 §4.2, §4.5; no within-interface baseline, §4.1), and
states the rule: results from different interfaces are never pooled or paired. Every run'schannelisapi
today; a manual channel, if added, must record interface, device and sign-in state. - A preregistered validation plan (
benchmark/PREREGISTRATION.md+ zh), written before any data and registered
by the first tagged release that contains it: does AIV readiness go with being cited (watch's citation rate) and
with absorption (supported_share.chars, computed for every site of a category from the savedspans; not
measured where a provider documents no positions)? Sites that compete for the same unbranded questions in 12 to 20
categories, drawn across the rubric's own band cuts; 8 asks per question per engine per wave (arXiv 2604.07585 §7),
two waves 12 weeks apart; every site of a category read on the same answers as its concurrent control (arXiv
2606.04362 §5); every site re-scored and logged each wave (arXiv 2506.11097 §6.4); a within-category permutation
test and four Holm-corrected secondary analyses written in advance (arXiv 2604.25707 §11.1), and what is published
whatever the result. It tests the instrument, not whether GEO works; a causal study would randomise page changes
(arXiv 2604.25707 appendix B). Linked from VALIDITY.md (+ zh) throughbenchmark/validity.py. No data collected yet.
Changed
- Evidence table: literature added to five rows, metadata only (
rubric/evidence-v1.1.md+ zh and the
evidenceobjects inrubric/v1.1.json).p2.question-intentadds arXiv 2603.25500 §5.2 Table 6, 2511.20867 §5.3
Table 5 and 2606.04362 §4: thin, slightly contrary.p2.freshnessadds 2607.15771 §2.5 Table 2-5: pages cited for
highly time-sensitive questions are newer than those cited for low-time-sensitivity ones, which describes the age
of cited pages, not an effect of showing a date.p2.sourced-statsadds the GEO paper's own all-sources result
(2311.09735 §5.2 Table 2), the rank-based counter-evidence (2506.11097 §6.2, §6.4) and 2402.11782.p3.listings
andp3.mentionsadd 2607.15771 §2.1 Table 2-1 and §3.2 Table 3-2, and state 2606.20065 §6.5 exactly (2.9% own
domain, 75.2% other companies). Those rows re-verified 2026-10-05. The README's limitations cite the external
evidence that two asks of one question differ: source overlap of 0.32 to 0.43 per industry with engines pooled and
0.23 to 0.51 per engine (2604.07585 §5 Tables 4 and 6), 0.66 for AI Overviews (2604.27790 §4.4 Table 4). No weight,
tier, score or evidence basis changed.
Fixed
- The README no longer quotes the GEO paper's "up to 40%" without its limits. It now says the gain was
measured in a simulated engine with one source rewritten at a time, that the top-ranked source lost 20 to 30%
when every source was rewritten (arXiv 2311.09735 §5.2 Table 2), and that a citation-rank benchmark found most
methods gave no significant gain (arXiv 2506.11097 §6.2).
Verify downloads: shasum -a 256 -c SHA256SUMS.
中文更新日志:https://github.com/jianruntech/geo-score/blob/v1.8.0/CHANGELOG.zh-CN.md
geo-score 1.7.1
Fixed
- A run started at one page no longer reads its site navigation as no links. The server-HTML links
observation (new in 1.7.0) kept only links inside the path the run was given, sogeo_score.py example.com/product
counted 0 links on a page with a full navigation bar and flagged it as navigation that appears only after
JavaScript runs. Links now count anywhere on the host, the start page itself aside, and the report names that page
as the page the run started from rather than the homepage ("links to N pages on this host"). Reported from a live
site the day 1.7.0 shipped. Never part of the score.
Verify downloads: shasum -a 256 -c SHA256SUMS.
中文更新日志:https://github.com/jianruntech/geo-score/blob/v1.7.1/CHANGELOG.zh-CN.md
geo-score 1.7.0
Added
engines[].timeout. Seconds to wait for one engine's answer, 10 to 600 (default 120), for a slow endpoint
or a model that searches and reasons at length. A timed-out request is still not retried.- "What the engines searched for" counts years and ranking words, and names searches that came back without
their queries. A line under the table gives, over every search the answers exposed, how many contain a year
(a four-digit number from 1900 to 2099 standing alone) and how many contain a ranking or comparison word:best,
top,vs,versus,review,reviews,compare,comparison,alternative,alternativesas whole words
in any letter case, or最好,最佳,推荐,排行,排名,哪家,哪个好,对比,评测anywhere. A search
counts once for each, and one that repeats the question's own words counts too: the line describes what was
searched, not what the engine added. When OpenAI, Gemini or Anthropic answers report a search but expose no query,
as an OpenAI-compatiblebase_urlthat returnsweb_search_callitems with noactiondoes, the section no
longer disappears: it says how many answers searched with no query text, per engine, andwatch asksays the same
under the answer. Perplexity and OpenRouter answers are never counted there, since their adapters record one
search per answer whatever happened. The summary gainssummary.searches.answers_searched_without_queries,
summary.searches.modifiers(total,with_year,with_ranking_word) andsearched_without_querieson each
row ofsummary.engines, and schema/watch.v1.json now describessummary.searches. A run saved before this
release gets them computed from its saved answers when it is reported; the run file is not rewritten. - Pages robots.txt blocks by path, and rules two readings disagree on, outside the score.
g.robotsreads only a
Disallow that closes the whole site, so a rule on one path left it green. Each run now holds the homepage, every
sampled page and the first 200 URLs on the audited host and path of the sitemapp1.sitemapread (the first child
of a sitemap index; when the homepage redirects to another host, that host's too, if robots.txt was answered from
there) against the rules each of the ten retrieval crawlers obeys, read as RFC 9309 section 2.2.2 says: the longest
matching pattern decides, an allow wins a tie,*and a final$are wildcards, and percent-encoding is compared
in one form. The report names each page a rule closes, the crawlers and the rule, and the pages where the
first-match reading of Python'surllib.robotparserand many older checkers (rules in file order, the first plain
prefix decides) says otherwise, with the rule each reading chose. Only the first 500 KiB of robots.txt are read, the
least RFC 9309 section 2.5 asks a crawler to parse, and the readings stop after 1,000,000 rule comparisons, counting
the URLs left as not checked, so a file of thousands of wildcard rules cannot hold a run up. Nothing is fetched for
it and it is never scored: rubric v1.1, every tier and the leaderboard are unchanged. - Whether pages are linked in the server HTML, outside the score. From the HTML the run already fetched (the
homepage, the scored sample and the section indexes sampling read), the report counts the<a href>links to pages
on the audited host and path, or on the host the homepage redirected to (apex to www), that a crawler running no
JavaScript can follow: links inside comments, scripts, styles and templates do not count, nor dojavascript:,
mailto:andtel:links. It says how many pages the homepage links to and how many of the other sampled pages
among them, flags a homepage that links to fewer than 3 ("navigation that appears only after JavaScript runs cannot
be followed by crawlers that do not run it"), and names the checked sitemap URLs no fetched page links to (the
homepage aside, since a crawl starts there). Only the fetched pages are searched, so a page they do not link to is
not shown to be orphaned, and the report says so. No request is added, and it is never scored. - hreflang annotations on the fetched pages, outside the score. On the same pages, the report reads each
<link rel="alternate" hreflang="…" href="…">in the server HTML (outside comments, scripts, styles and templates; a
feed's alternate left out) and eachrel="alternate"link with anhreflangparameter in the page's HTTPLink
header, everyLinkline read (a response's repeatedLinklines used to keep only the last). It says how many
pages carry annotations, how many there are and how many came in aLinkheader, and lists, as found: pages that
carry annotations after the end of the head (after</head>or the<body>tag), since Google's hreflang
documentation says the<link>tags must be inside a well-formed<head>(they are still counted with the others);
pages whose annotations leave the page itself out; codes outside the subset of BCP 47 that Google's hreflang
documentation supports (an ISO 639-1 language, optionally an ISO 15924 script and an officially assigned ISO 3166-1
alpha-2 region, orx-default), soen_US,jp,en-UK,es-419and the registry's six deprecated two-letter
language subtags (iwand the like) are listed, each with why; a code that names more than one URL on one page; URLs
named in annotations that are not absolute, as written; and URLs on another host, which hreflang allows and the
report only notes. A>inside a quoted attribute value does not end a<link>, a backslash-escaped quote in a
Linkheader parameter does not end it, and a default port or a trailing dot on a host names the same host. Return
links are checked only between the pages the run fetched: a pair where one page names the other and is not named
back is listed, and the other URLs the annotations name are counted as not checked. A site with none gets one line.
The codes are the 184 two-letter language subtags the IANA Language Subtag Registry (File-Date 2026-09-17) does not
mark deprecated and the 249 officially assigned ISO 3166-1 alpha-2 codes among its region subtags; the registry and
Google's page (last updated 2026-09-21) were read on 2026-10-04. Only the first 1,000 annotations of a page are read, and the rest are counted. No request is added, and
it is never scored: rubric v1.1, every tier and the leaderboard are unchanged. - The text report, the MCP
score_sitesummary and the Action's job summary show all three under "Outside the
score". The JSON report carries them in a new optionalaccessfield (access.robots_paths,
access.server_links,access.hreflang), registered in schema/report.v2.json; each lists the first 50 pages, and
its counts cover them all. The terminal report quotes a page's path, a robots.txt pattern and an hreflang code or
href as«site text: …», the Markdown views as code. A new e2e fixture,hidden-pages, has a path rule that closes
its pricing page, rules the two readings part on, and a menu only a script writes; another,hreflang, has a page
that is not named back, annotations in twoLinkheader lines, and the code and href slips above; thehostile
fixture carries its injection text in an hreflang code and href. Every fixture'sexpected.jsonpins the new
counts. - The Claude Code skill installs with
npx skills add jianruntech/geo-score(the open skills CLI;-gfor a
global install). Checked on 2026-10-04: it finds one skill,geo-score, and copies the whole repository, about
9 MB, into.claude/skills/geo-score. The README'sgit clonestays as the way without Node.js.
Changed
- A ledger with no priced answer says the cost is unknown. It read
$0 + 4 answers with no price (add prices to geo-score-watch.json), which looks like a free run; it now readscost unknown: 4 answers with no price (add prices to geo-score-watch.json). When some answers are priced it still reads$0.07 + 12 answers with no price.
The MCPruntool's message for a cancelled run uses the same words. The JSON ledger is unchanged. - The ledger says what failed. When calls failed,
watch reportadds a line with how many, and how many HTTP
requests were sent in all, since a provider may bill a request it timed out on. A path outside the current
directory is printed absolute instead of as a chain of../.
Fixed
- OpenAI engines send the question as a one-message list. The Responses API accepts a string or a list;
some OpenAI-compatible endpoints that abase_urlpoints at accept only the list and refused every call.
Official OpenAI answers the same either way. - The crawler reference no longer says that blocking an opt-out token costs nothing.
reference/ai-crawlers.md
put the cost of blocking an opt-out token at "nothing, in citation terms" and closed by saying that blocking every
training crawler and both opt-out tokens "costs nothing". Google's common crawlers page (last updated 2026-07-14,
checked 2026-10-04) saysGoogle-Extendedgoverns grounding in Gemini Apps and in Grounding with Google Search on
Vertex AI as well as the training of Gemini models, and does not affect inclusion or ranking in Google Search. The
reference now says what each opt-out token and training crawler controls, citing the vendor's page with its date:
Applebot-Extendedgoverns training only, and Apple's AI answers are opted out of withnosnippet(About Applebot,
published 2026-09-04); Meta's crawler page (undated, checked 2026-10-04) describesMeta-ExternalAgentas crawling
for training and also for indexing content to improve Meta's products, so it is no longer implied to be
training-only;anthropic-aiis not in Anthropic's current crawler article, so whether it sends requests is marked
no data instead of falling under a heading that said opt-out tokens se...
geo-score 1.6.0
Added
watch reportshows who is named for each question. A new table, "Who is named for each question",
lists every question with its answers and, out of those, how many cite and name you, how many cite and name
each competitor, how many name no tracked brand, how many searched the web (yes, no or unknown), and the
tracked brand cited most often, given only when the question has at least 4 answers and one brand is ahead of
every other. Every value is a count. Whether an answer searched is derived from what every run already saves
(a reported search, an exposed search query or a citation); Perplexity and OpenRouter answers are always
unknown, because their adapters record one search per answer whatever happened. Three lines under the table
say what the counts cannot: an answer given without a search cannot cite anything, "no tracked brand named"
covers only the brands in your config, and why some answers are unknown. The summary gains
summary.by_question($defs/by_questionin schema/watch.v1.json), the JSON report and the MCPreporttool
carry it, and it is an optional property of the tool's output schema, so the MCP contract only grows. A run
saved before this release gets it computed from its saved answers when it is reported; the run file is not
rewritten.- Suggested question types, and
watch checkcounts them.intent_tierstays free text; cli/README.md and
guide/watch-methodology.md now suggest eight labels, each with one line and an example:list,vs,
alternative,pricing,problem(a situation that names no category or brand),persona(who it is for),
verdict(is it worth it, can it be trusted) andchannel(franchise, wholesale, distributor, procurement).
watch checkprints how many questions have each type, with a note, never a failure, when every question has
the same type or none has one.watch initwrites two more placeholder questions, aproblemrow and a
verdictrow. target_urlis read:watch reportcounts how often the page a question tracks is cited. The column was
documented and saved but nothing used it. It is now resolved when the questions load: an absolutehttp(s)URL
as written (fragment andutm_*parameters dropped), or a path starting with/joined to the brand's first
domain.watch check,watch runand the MCPrun(its dry run too) andstatustools warn about any other
value, which the run ignores, and about a page that is not on the brand's domains or URL prefixes, which is still
matched. Each answer recordstarget_cited(one of its
citations is that page) andother_own_cited(it cites another URL of the brand's, not that page), both null
when the question tracks no page. The match is the exact page oncewww, letter case, a trailing slash, the
query string, the fragment and percent-encoding are set aside;/en/xis not/x, and a citation left as an
unresolved redirect counts as neither. A new table, "Pages you track", gives per page its questions, their
answers, how many cite the page, how many cite another of your URLs instead, and the URL of yours those
answers cite most, exactly as cited. The CSV export gains a last column,target_cited(Y,Nor empty).
Schema:target_citedandother_own_citedin$defs/sample,summary.tracked_pageswith
$defs/tracked_page, and a description of the savedtarget_url. The MCPreporttool's output schema gains
the optionaltracked_pages, so the MCP contract only grows. A run saved before this release gets the table
computed from its saved citations when it is reported.watch reportshows who is named first. Each answer records, for the brand and each competitor,
mention_order(1 for the brand named before every other brand in the config, by the earliest match of its name,
aliases or domains with the matcher ofmentioned; null when it is not named, or when two brands are first named
at the same place) andmention_share(the share of the answer's sentences that name it, every sentence counted
the same; a sentence ends at a line break, at。!?, or at.!?before whitespace). With competitors in the
config, "By engine" gains a column, "named first (tracked)": the answers that name you first among the brands in
the config, out of the answers that name at least one of them. "Share of voice" gains "named first (tracked)" and
"sentence share", the meanmention_shareover the answers that name each brand, with three lines on how they are
read. Brands the config does not list are not seen, so first among the brands you track may not be first in the
answer. Both are read from the answer as saved, which keeps its first 20,000 characters. Schema:mention_order
andmention_sharein$defs/sampleand its competitors, andnamed_first,any_namedandmention_shareon
the rows ofsummary.enginesandsummary.share_of_voice. The MCP tools' schemas are unchanged. A run saved
before this release gets them computed from its saved answers when it is reported.watch run --dry-runsays what the plan can detect, andwatch reportcan pool runs. The dry run prints the
plan's resolution: with E engineswatch difftests E + 1 rows (one row with one engine), and it can call a
change only when at least k questions move the same way between two runs, k being the smallest n with rows × 2 /
2^n < 0.05, the screen diff already applies: 6 questions with one engine, 7 with two, 8 with three to five. A set
with fewer questions gets "this plan cannot detect a change": more questions help, more engines raise k, and more
asks per question do not lower it. The MCPruntool's dry run prints the same lines.watch report --since DAYS
(or--runs A..B) counts the saved runs of a period as one set of answers, when each asked the same questions
(the samequestions_hash) as many times each (the sameruns, so every run weighs the same) on the same engines
with the same models, from the samecountry, with the same brand domains and URL prefixes as the newest;
otherwise it exits 1 and names every run that differs and why. It gives the runs, the calendar days they cover,
the answers, the headline rate and a row per engine, with a 95% interval that resamples questions, each with every
answer it got in every run, and the run report's notes (the API channel, questions that name the brand, what
counts as not measured, and that it does not predict traffic, rankings or revenue). There is no pooled diff.
Schema:$defs/pooledin schema/watch.v1.json. MCP: thereporttool gains an optionalsinceargument and its
output schema the optionalrunsanddays, so the MCP contract only grows. guide/watch-methodology.md gains
section 8, "One ask, one day, one month".
Changed
- Install from PyPI. geo-score has been on PyPI since 1.5.0, as the same wheel and sdist the GitHub release
attaches, and the install lines now lead with it:uvx geo-score@X.Y.Z example.com,
pipx install geo-score==X.Y.Z, anduvx geo-score@X.Y.Z mcpfor MCP clients (the install buttons, the Claude
Code line, the standard config, every client in guide/mcp.md and the Claude Desktop bundle's watch-init hint). The
release wheel URL stays as the line for a machine without PyPI, the git tag as the from-source line, and thecurl
line as the way to run level 1 with nothing installed. server.json lists the pypi package beside the.mcpb, so
the MCP Registry offers it from this release on, and README.md, the package's description on PyPI, carries the
registry'smcp-nameline (1.5.0's description on PyPI has none, so a 1.5.0 listing names only the.mcpb).
release.py --bumpmoves everygeo-score@X.Y.Zandgeo-score==X.Y.Z, and--executeends with the PyPI
upload (--no-pypileaves it out,--publish-pypidoes it alone): it uploads exactly the wheel and sdist that
SHA256SUMSnames, with twine pinned inTWINE_VERSION, then checks PyPI's SHA-256 of each (RELEASING.md,
step 8). Its requests to GitHub and PyPI are tried three times, and a stop that a rerun can get past names the
command.
The messagewatch,mcpand--askprint whengeo_watch.pyis missing (the piped one-line run) names the
package first too, then the wheel.
Fixed
watch diffcompares a question only with itself. It paired questions byquery_idalone, so a question
reworded under the same id was compared with its old wording, and the edit could read as a change (and make
--fail-on-dropexit 4). The question text is now compared case-folded, with whitespace collapsed: a question
whose text changed is left out of every row and named in the notes (at most 10 ids, then "and N more"). The
6-question minimum and the screen for rows no test could call apply to the questions left, so on runs with
edited questions a verdict can change, and can becometoo few questions to tell. The notes also say when the
runs asked from different locations (country) and when an engine id moved to another provider. Run files gain
an optional top-levelquestions_hash(SHA-256 over the sorted pairs of id and normalised text, defined in
schema/watch.v1.json), andwatch runsshows its first 8 characters; for older runs it is computed from the
saved questions.- An API that changes its answer format no longer reads, without a word, as an engine that stopped citing.
Each adapter reads one format per provider, so a renamed citation field gave answers with no citations at all,
andwatch diff --fail-on-dropcould exit 4 on what was a format change.watch reportnow warns under "By
engine" when an engine answered 5 or more questions and none of its answers carried a citation of any domain.
watch diffadds awarningto an engine's row, the same words to the notes and to stderr, when the share of
its answers carrying a...
geo-score 1.5.0
A release about reading sites the way the rubric says, and about failing CI for the right reason. The two
crawler gates now read only the ten retrieval crawlers, choose each crawler's robots.txt group as RFC 9309
does and retry a request that got no answer before they count it; a bot-challenge page is no longer scored
as a page that needs JavaScript; every page is read in its own language; and a figure counts as sourced
only when its source sits beside it. These fixes move scores, and each one below names its check and its
direction. For CI there are --fail-on-gate and --assert (a gate at zero caps a score at 40, and a site
that would otherwise score 40 or more reads exactly 40, which --fail-under 40 lets through),
geo-score diff and --baseline to compare two reports check by check, and a shields.io endpoint badge.
A site that cannot be scored fails fast and says why in one line, in JSON under --json, and a scoring
run can no longer hang. Around the CLI: a schema and typo warnings for the watch config, installs from the
release wheel without git, a gate that pins its tools, CI examples beyond GitHub, forms for disputing a
leaderboard row, and a benchmark that can be re-measured and taken apart. The leaderboard itself is still
the 1.3.0 run of 2026-09-27.
Upgrading from 1.4.0. Level-1 scores move only where 1.4.0 read a site wrong, and every such fix is
under Fixed with its check and direction: the gates (g.robots, g.reachable, and g.ssr for sites
behind a bot challenge, with Japanese or Korean pages, or with English pages that carry more than 20 Han
characters), the checks that read page text (p2.answer-passages, p2.question-intent,
p2.sourced-stats, p2.named-author, p4.answer-shape, p4.cn-engines and p3.knowledge-graph), and
the checks whose file, logo or crawler request is now retried (p1.llms-txt, b.llms-full, b.ai-txt,
p1.organization and p4.cn-engines). Three invocations exit differently: a site that serves only
bot-challenge pages exits 2 instead of being scored; an unexpected error exits 2 in one line instead of
a traceback that exited 1; and --sample below 1 exits 2, where 1.4.0 crashed on 0 and scored fewer
pages on a negative value. diff is now a subcommand, so a host named diff is written
https://diff. audience_language can hold languages other than en and zh-CN, and the report gains
only the optional elapsed_s.
Added
- A glossary and a Chinese CLI reference.
guide/glossary.md(andguide/glossary.zh-CN.md) defines the 49
terms the README, the reports and the guides use, each with where it appears and which page goes deeper;
cli/README.zh-CN.mdis the CLI reference in Simplified Chinese, with a language switch at the top of both. --fail-on-gateand--assert EXPR, in the CLI and as the Action inputsfail-on-gateand
assert. A gate at zero caps the score at 40: a site that would otherwise score 40 or more reads
exactly 40 and passes--fail-under 40(fail-under: 40).--fail-on-gatefails any capped site: it
exits 1 when anyg.*check scores 0, and names the gate and its evidence; it is the same as
--assert 'g.*>0'.--assertholds one check, or an fnmatch glob of checks, to a value, e.g.
g.robots=maxorp1.llms-txt>=4, and is repeatable. A check the run could not observe is skipped,
never failed. A malformed expression, or a glob that names no check, exits 2.
Both are opt-in and change no score, and their results go to stderr, never into the report.geo-score diff BEFORE AFTER [--format text|md|json] [--fail-on-drop]compares two level-1 reports
check by check. Each check shows its points before and after, its max, and a state:up,down,
same,became unobservable,became observable, orunobservablewhen neither report observed it.
The comparison also shows the score, band and gate cap before and after, names each gate that reached
or left 0, and lists the pages added to or removed from the sample. When the two samples differ, it
says so, reads the totals as ±5 (benchmark/REPRODUCIBILITY.md) and namesg.reachableas the least
stable check. A check that became unobservable is never shown as a drop to 0.--format jsonfollows
the newschema/diff.v1.json. Two reports of different rubric versions exit 2, and reports from
different releases, or of two different targets, are compared with a warning on stderr.diffneeds
onlygeo_score.py.--baseline REPORTon a scoring run scores the baseline's pages again (as--urls-fromdoes) and
prints the same comparison after the report. It cannot be combined with--urlsor--urls-from.
Under--json, stdout carries only the report and the comparison is not printed: write the report
with--json-outand rungeo-score diffon the two files to see it.--fail-on-dropstill exits 4
and names each drop on stderr.--fail-on-drop(with--baseline, or ondiff) exits 4 when a check lost points, a gate reached
0 or the band fell, and prints one stderr line per reason. It is opt-in, aswatch diff --fail-on-drop
is, so no existing invocation changes its exit code. A failed--fail-under,--fail-on-gateor
--assertstill exits 1.- Action: a
baselineinput,deltaanddropped-checksoutputs, and a "Changes since baseline"
section in the job summary, with the check ids read from the baseline escaped. The baseline may be the
same file asjson-out. - A badge that stays current.
--badge-json FILEwrites the badge as shields.io
endpoint JSON:
{"schemaVersion": 1, "label": "AIV readiness", "message": "71/100 Solid", "color": "#2f8f52", "cacheSeconds": 86400}.
The SVG and the JSON take their message and colour from one helper, so the two cannot disagree. Host
the file anywhere andhttps://img.shields.io/endpoint?url=<its URL>draws the badge. The Action gets
abadgeinput: a path ending in.jsongets the endpoint JSON, and any other path gets the SVG. A
badge-pathoutput says where the file went. - A JSON error instead of nothing. Under
--json, every exit 2, usage errors included, prints
{"error": {"kind", "message", "target"}, "tool": "geo-score-cli/X.Y.Z"}on stdout, where stdout used
to be empty;diff --format jsondoes the same. The newschema/error.v1.jsonnames the kinds:dns,
refused,tls,timeout,droppedandno_responsewhen nothing answered;http_<code>,
no_content,slowandbot_challengewhen the site answered;deadlinewhen the run ran out of
time;not_publicunder the MCP server's address guard; andusage,rubric_mismatch(reports of two
rubric versions given todiffor--baseline),errorandinternal. A minor release may add
kinds; read an unknown one aserror. --timeout SECONDS(default 15, as before; up to 300): how long each request may wait for the
server before it counts as unanswered.elapsed_sin the JSON report (optional,schema/report.v2.json): how many seconds the run took
on the machine that ran it. The footer prints it too (· took 23.4 s). It describes the run, not the
site, anddiffignores it.- A schema for
geo-score-watch.json, and a warning for every key nothing reads.
schema/watch-config.v1.json(JSON Schema draft 2020-12) describes the file, andwatch initnow
writes"$schema"as its first key, so an editor checks and completes it. A mistyped key used to
vanish without a word:budgetforbudget_usdmeant no budget, andmax_callformax_callsthe
default cap of 120.watch checknames every unknown key at every level (top,brand,
competitors[i],engines[i],prices.<id>), with the nearest known key when one is close, as in
unknown key 'budget' (did you mean 'budget_usd'?); it is ignored, and flags apricesentry that
names no engine.watch runand the other commands that read the config print the same warnings on
stderr; the MCPstatustool lists them, andrunshows them with its plan and after a saved run. A
key named like a credential (api_key,OPENAI_KEY,api_token,client_secret,credentials…)
is pointed at environment variables instead, and its value is never printed. Warnings only: no exit
code changes. - CI beyond GitHub.
examples/ci/gitlab-ci.ymlis a GitLab job onpython:3.12-slim: it downloads
geo_score.pyfrom a pinned release, checks it againstSHA256SUMS, scores once, keeps
geo-score.jsonas an artifact and handsgeo-score.xmlto GitLab as a JUnit report, so a check at
zero shows as a failed test.examples/ci/README.mdgives the one directive that Bitbucket Pipelines,
Azure Pipelines, Jenkins and CircleCI each need to show the--junitfile. - Where to report what. A Benchmark correction form disputes a leaderboard row's score, gate
reason, label or not-scored reason, with a fresh--json-outreport and the tool version; a Tool
problem form covers the CLI, the Action, the MCP server, pip/pipx/uvx, the MCPB bundle, the Claude
Code plugin, the Gemini CLI extension and Docker. The issue chooser links the troubleshooting guide and
the MCP troubleshooting section, and a newSUPPORT.mdsays which form takes what. The forms applied
benchmark,engineandrubric, labels the repository did not have, so those issues arrived
unlabelled:.github/labels.jsonnow lists every label, andscripts/sync_labels.pycarries the list
to GitHub (a dry run by default;--applycreates and updates,--checkexits 1 when a label is
missing, and it never deletes one). It prints a proxy's host and port, never the credentials in
HTTPS_PROXY. - Benchmark: a re-runnable retest study.
benchmark/retest.pytakes two or more run directories
(results.json plus raw/) and records each run's tool, date, Python, declared vantage and pinning....
geo-score 1.4.0
The MCP server becomes a first-class way to use geo-score: its own README section with install
buttons, a guide for every client, one-click packages, both protocol eras, and a surface that says
what costs money and lets you turn that off.
Added
- MCP server, both protocol eras. 2026-07-28 (
server/discover, per-request_meta, cache hints)
alongside 2025-11-25, 2025-06-18, 2025-03-26 and 2024-11-05 overinitialize. An unknown legacy
version gets 2025-11-25; an unsupported modern one gets-32022with the supported list. Batches are
accepted only on 2025-03-26, and acursoris an error because the server never issues one. - MCP tools say what they cost.
score_siteis listed first. Every tool has a title and all four
annotation hints, descriptions open with the cost ("Free, no API key." / "Paid: …"), input schemas
reject unknown keys, and nullable fields useanyOf, which strict clients accept.askandrun
now return structured results, and arundry run returns its plan as data. - Cancelling stops spending.
notifications/cancelledstopsaskandrunbefore their next
call; a cancelled run is not saved, and its calls and cost go to stderr. Tool calls run on worker
threads (up to 4), so the server still answerspingduring a long run. At end of input it stops the
paid calls at once and lets free ones answer. Onerunat a time: a second is refused while one is in
progress, so parallel calls cannot double the spend of one approval. - Robustness. A malformed run file or config is an error answer, never the end of the server;
arguments that are not an object are refused rather than read as{}; stdio is UTF-8 on every
platform, and a stream that cannot carry a character gets it escaped instead of losing the answer;
statusnever shows a value pasted intokey_envand reportsrunas ready only when it could start;
GEO_SCORE_MCP_EXCLUDE_TOOLSand--exclude-toolsadd up. - Choose the tools.
--toolsets score,ask,watch,--tools,--exclude-toolsand--read-only,
each with aGEO_SCORE_MCP_*environment twin.--read-onlygives a server that cannot spend. A new
read-onlystatustool reports the version, the config it found, which engines have keys (never the
keys) and the caps. - Prompts and completions.
audit_site,check_citations,weekly_watch,compare_runsand
explain_check, with completion for check ids and run ids, and ageo-score://runs/{run_id}resource.
None of them contains fix instructions: measurement is public, remediation is not. - Packages. Every release attaches a Claude Desktop bundle (
geo-score-X.Y.Z.mcpb) and a
server.jsonfor the MCP Registry, both covered bySHA256SUMS. The repository is also a Claude Code
plugin marketplace (/plugin marketplace add jianruntech/geo-score) and a Gemini CLI extension, and
has a Dockerfile for the stdio server. Nothing is published to a registry, PyPI or a container
registry yet. - Docs. A
## MCP serversection in both READMEs with install buttons (Cursor, VS Code, VS Code
Insiders, LM Studio), a standard config, the tool table and example prompts;guide/mcp.mdwith the
exact setup for Claude Code, Claude Desktop, Codex, Cursor, VS Code, Gemini CLI, Devin Desktop and
Zed, plus timeouts, security and troubleshooting.scripts/mcp_docs.pygenerates the tables and
buttons from the server, andcheck.shfails when they are stale. - Checks.
scripts/check_mcp.pydrives the server over raw stdio in both eras on every tested
Python;check.sh --liverepeats it through the official MCP Inspector (2.8.0,--strict) and calls
score_site.scripts/test_packaging.pychecks every manifest against the code. A snapshot and a
1.3.0 contract baseline stop a tool or a required argument from changing by accident. - SECURITY.md has a per-tool table of network access, files and spending for the MCP server.
Changed
- The leaderboard is re-run with 1.3.0 (2026-09-27): 317 of 370 sites scored, median 56 (95% CI
52–58), range 11–99; 83 sites (26%) have a gate at zero (47 need JavaScript, 23 block AI crawlers,
13 serve an error). Across the 312 sites scored in both runs the mean change is −0.4 points; 72% move
by 5 points or less, which mixes the 1.3.0 scoring fixes with 18 days of site changes and is not a
reproducibility figure (that study isbenchmark/REPRODUCIBILITY.md, same tool, two runs). - Every site's full report from the run is in
benchmark/raw/, andvalidate.pychecks it against
results.json(rule 29). - The benchmark README's findings are computed rather than written by hand: the headline share,
the top gaps, and which checks drive the Chinese-market gap. That gap is 20 points (95% CI 16–28),
and its two largest per-check drivers are heuristics the CLI also read lower than a human on the
Chinese site in the hand-audit comparison; the README says so. - The READMEs link the leaderboard on www.jianruntech.com/leaderboard, which is kept current; the
GitHub Pages copy cannot be rebuilt while Actions and Pages are disabled for the organisation.
Verify downloads: shasum -a 256 -c SHA256SUMS.
中文更新日志:https://github.com/jianruntech/geo-score/blob/v1.4.0/CHANGELOG.zh-CN.md
geo-score 1.3.0
A release about trust in the measurement: outputs CI systems read natively, a pinned sample for
before/after comparisons, change detection that no longer calls noise a change, offline end-to-end
tests on twelve fixture sites, a measured CLI-vs-hand-audit agreement, and an evidence basis for every
check. The fixtures also found places where the CLI read sites against the rubric (grouped
robots.txt rules, single-page apps' catch-all pages, empty markup, stale dateModified, bylines,
Chinese dates and figures); those are fixed and listed below with their direction. The percentage can
no longer pass 100, and the MCP server now keeps every connection on public addresses.
Upgrading from 1.2.0. Level 1 scores move only where a check read a site wrongly; each fix and
its direction is listed under Fixed, and the public leaderboard (measured with 1.1.0) is not
re-scored. In watch, the headline rate leaves out questions that name the brand, intervals widen
when a question has several answers, and some diff verdicts that said change now say
within noise or too few questions to tell; the 1.2.0 fields keep their meaning. Evidence strings
now quote site text inside «site text: …». The CSV report has an eleventh column, branded.
Added
- One run, several outputs:
--json-out FILE,--sarif FILE(SARIF 2.1.0: one rule per check,
one result per check below full marks) and--junit FILE(one test per check; a check at zero
fails, an unobservable one is skipped). All are written from the same fetch as the report on
screen, so they cannot disagree with it. Single site only. - Install as a command:
pyproject.toml, sopipx install git+https://github.com/jianruntech/geo-score
anduvx --from git+… geo-scorework. Installsgeo-scoreandgeo-score-mcp(the MCP server over
stdio). Still no dependencies. STABILITY.md: what is a public contract (check ids, score arithmetic, the JSON schemas, CLI
flags, exit codes, Action inputs and outputs, MCP tool names, SARIF rule ids), what semver means
here, and that nothing is sent anywhere (no telemetry).- Action inputs
sarif,junitandannotations; outputsgate-cappedandreport-path. - Pin the sample:
--urls FILEscores exactly the given pages (up to 8, all on the scored site);
--urls-from REPORT.jsonre-reads the pages an earlier report sampled. Sampling is the main source
of run-to-run variance, so a before/after comparison on pinned pages shows what changed on the
site. The report sayssample_source: "pinned". Action inputurls-from. - Per-check documentation links: every check in the JSON report carries
doc_url, its definition
in the rubric pinned to the release tag, so the link never drifts from the scorer. SARIF rules link
there too, and--explainprints it under each check below full marks. Optional field in
schema/report.v2.json. - Troubleshooting, in English and Chinese: score movement, checks shown
as—, the gate cap, exit code 2, proxies, redirects, the leaderboard's tool version,watch
exit 3, Action annotation limits, CJK thresholds. The terminal report links to it. - Intervals on the leaderboard's medians.
benchmark/stats.pycomputes a seeded bootstrap 95%
interval for the overall median, each sector median and the Chinese-market gap;benchmark/README.md
and both READMEs quote them: overall 56 (53–58), the gap 19 points (12–26). The hosted page shows
medians without intervals until it is rebuilt and redeployed. Sectors with fewer than 10 scored sites (China / cloud and Professional services, 9
each) no longer print a median;validate.pycounts scored sites, not listed ones.
benchmark/run.py --readme-onlyregenerates the README without re-scoring. - The leaderboard can move host:
GEO_SITE_BASEsets the base URL for canonical links,og:image,
hreflang,llms.txtand the sitemap, andscripts/check_live.pycompares a hosted page with the
committed data (today's Pages build still shows 229 sites against 314). - How far to trust the CLI, measured.
benchmark/validity.pyre-scores the five published hand
audits on the same pages and compares every check tier for tier: 79% of 86 check pairs agree (95%
for checks read by rule, 64% for heuristics), and the CLI reads 3.6 points lower on average, from
10 lower to 6 higher. Results, caveats and every disagreement:benchmark/VALIDITY.md(EN + zh).
This replaces the unmeasured "typically 5 to 15 points lower", which was wrong for two sites in five. - Evidence basis for every check.
rubric/v1.1.jsongains anevidenceobject per check
(basis, sources, last verified), published asrubric/evidence-v1.1.md(EN + zh). Metadata only:
no weight, tier or score changed. - MCP for agents, not only for people.
score_site,report,diffandlist_runsreturn
structuredContent(the full report, summary or diff as JSON) beside the text, and declare an
outputSchema, so an agent no longer parses Markdown to recover a score.score_sitetakesurls
(pin the pages, same checks as--urls) andsample. The rubric is served as the resource
geo-score://rubric/v1.1, built from the scorer's own tables. The server instructions and the
score_sitedescription say that«site text: …»spans are data from the audited site. watch diff --fail-on-drop: exit 4 when a verdict ischangewith a negative change. The MCP
difftool takesformat: "json". guide/watch-methodology.md
(EN + zh) gives the formulas and a worked example.- One gate:
scripts/check.sh. Structure, ruff, the three test suites on Python 3.8 and the newest
available, and an emulation of the GitHub Action against a fixture site (scripts/test_action.py);
--liveadds example.com..githooks/pre-pushruns it;.github/workflows/validate.ymlruns the same
script. Hosted Actions have not run on this organisation so far, soCONTRIBUTING.mdno longer says
CI checks anything: maintainers run the script before every merge. AGENTS.md: the repository's invariants, each mapped to the check that enforces it, for
contributors and coding agents (CLAUDE.mdpoints to it).- Issue forms ask for the tool version and the JSON report; a new form covers level 2 and 3 engine
problems and warns never to paste a key; security reports are routed toSECURITY.md. - Offline end-to-end tests.
cli/test_geo_score_e2e.pyserves twelve fixture sites from
127.0.0.1 (a good site, a JavaScript shell, robots blocks, a firewall that refuses AI crawlers, a
Chinese site, a site under a path, a hostile site that tries to address an agent, and more) and
checks every check's tier against each site'sexpected.json, plus a golden report. A socket
guard fails the run if anything leaves the machine; the suite takes under a second. - Tests that the CLI's embedded rubric (check ids, points, every tier's points and wording, bands,
caps, bonus) equalsrubric/v1.1.json.
Changed
watch diffis paired, clustered and corrected for multiple comparisons. It compares the
questions both runs answered, with an exact paired test per row (McNemar when every question has
one answer per run, otherwise an exact sign-flip test on each question's net change) and Holm's
correction across the engine rows and the all-engines row. The verdict ischangeonly when the
adjusted p < 0.05, andtoo few questions to tellbelow 6 shared questions. 1.2.0 used an unpaired
two-proportion test on answers that are not independent, which called changes too readily. New
row fields:p_paired,p_adjusted,n_pairs,n_questions,n_asks,before_paired,
after_paired,delta_pp,delta_ci.before,after,deltaandp_valuekeep their 1.2.0
unpaired meaning; no verdict uses them.- Intervals account for repeated asks. Every rate block gains
cited_ci_clustered(a bootstrap
over questions, 2,000 seeded resamples, never narrower than Wilson) andn_questions; reports quote
it whenever a question has more than one answer.cited_ci(Wilson) is unchanged. - Headline rates leave out questions that name the brand. Questions are flagged at run time
(brand_in_query,competitor_in_query) with the matcher used on answers; the headline covers the
other questions and branded ones get their own line and n. An optionalbrandedcolumn in
queries.csv(auto,yes,no) overrides the flag,watch checkwarns when more than 30% of the
questions name the brand, and the CSV report gains a trailingbrandedcolumn. - The Action scores the site once. It used to fetch and score twice (once for JSON, once for
the log), which doubled the time and could print two different scores. The job summary, the
annotations and every file now come from one run. - The Action annotates checks at zero: gates as errors, other checks as warnings (GitHub shows up
to 10 of each; the rest are in the job summary). geo-score watch --helpnames the command you typed (geo-scoreafter an install).--comparewith three or more sites no longer says "Both scored".- When a site redirects to another domain, the full report says so and the JSON carries
landed
(optional), the URL that was scored.
Fixed
- Scoring fixes the offline fixtures found. Each moves a score only where the old reading
contradicted the rubric:g.robots: severalUser-agentlines share the rules that follow them (RFC 9309 2.1). The rules
were attached to the last line only, so a group listing GPTBot, ClaudeBot and CCBot over
Disallow: /read as blocking only CCBot. Down, to 0 and the gate cap, where grouped rules block
retrieval crawlers; up (3 to 5) where a groupedAllownow counts as an explicit permission.p1.llms-txt,b.llms-full,b.ai-txt,p1.sitemap: a single-page app that answers every path
...
v1.2.0 — levels 2 and 3: does AI actually cite you?
Added — levels 2 and 3: does AI actually cite you?
geo-watch 0.1.0 is merged into this repository. One tool now measures both halves of AI
visibility: whether engines can cite a site (the score, free) and whether they do
(with the user's own API keys). The citation half is never summed into the score, as the
rubric has always required.
- Level 2,
--ask "question"(repeatable, up to 5): after the readiness report, asks
ChatGPT, Perplexity, Gemini and Claude the question through their search-enabled APIs,
one ask per engine that has a key, and fills the report'scitationobject. Without a key
the block says not measured and the score is unchanged. - Level 3,
geo_score.py watch: a fixed question list asked every week; citation rate per
engine with 95% Wilson intervals, share of voice against competitors, the third-party
sources the engines cite, the searches they ran, a cost ledger with caps (max_calls
is exact;budget_usdcan be exceeded by at most one call per engine), anddiffwith a
two-proportion test. Runs followschema/watch.v1.json. geo_score.py mcp: a dependency-free stdio MCP server withscore_site(level 1) and
ask,run,list_runs,report,diff(levels 2 and 3).schema/report.v2.json: the reservedcitationobject gains optionalchannel,
answered,mentioned_count,per_engine,top_sourcesandcost_usd. Additive; every
1.1.x report still validates.examples/ci/watch-weekly.yml: the weekly schedule.
Levels 2 and 3 need cli/geo_watch.py from the same release next to cli/geo_score.py. The
one-line curl | python3 install still runs level 1 on its own. Scores are unchanged from 1.1.2.
For geo-watch 0.1.0 users: rename geo-watch.json to geo-score-watch.json and run
python3 cli/geo_score.py watch … instead of geo_watch.py …. Runs now live in
.geo-score/watch/runs/ with format geo-score/watch.v1; 0.1.0 run files are not read. The CSV
export renames competitors_cited to other_domains_cited and drops the always-empty
wrong_fact column.
Full changelog: CHANGELOG.md · 中文:CHANGELOG.zh-CN.md
v1.1.2 — a network that cannot reach something is not evidence it is missing
Changed — scores move
A site scored 88 and then 95 on two consecutive runs from one machine, and another
scored 71, 40, 71 — both outside the ±5 this repository promises. Three causes, one
principle: a network that cannot reach something is not evidence that it is missing.
- p3.sameas counted every non-200 as a dead profile. Timeouts, resets and 451 (LinkedIn,
X, YouTube, GitHub and Wikipedia are unreachable from mainland China) and bot walls
(401 / 403 / 429 / 999 from Crunchbase, Zhihu, Facebook, LinkedIn) are now left out;
only 404 / 410 count as gone. When no link can be checked, the check leaves the
denominator instead of scoring 0. The same 12 profiles resolved 1 of 12 from a
mainland network and 9 of 12 from overseas. - p3.knowledge-graph scored 0 when some Wikidata / Wikipedia lookups timed out and
the rest found nothing. Absence now needs every lookup to answer; a partial answer is
unobservable. (All lookups failing was already unobservable.) - Sampling fetched the sitemap once. One dropped request sent discovery to homepage
links, a different eight pages were scored, and the middle run fell under the g.ssr
gate (capped at 40). The homepage, robots.txt, sitemap, section indexes and sampled
pages are now retried on timeouts, 429 and 5xx. Reports carrysample_source.
After the change, three runs each: jianruntech.com 95 / 95 / 95 (direct) and 93 / 93 / 95
(overseas); stripe.com 75 / 75 / 75 (direct) and 70 / 72 / 70 (overseas, concurrent).
The published benchmark was measured on 2026-09-09 with the earlier logic and has not been
re-run. A site whose sameAs or knowledge-graph lookups hit an unreachable host can score a
few points higher today than its leaderboard row. Reports now say geo-score-cli/1.1.2, so the
two can be told apart.
Fixed
- gzip bomb. The 4 MB read cap applied to the compressed stream, and
gzip.decompress
has no output limit, so a few MB of crafted gzip could inflate to gigabytes and kill the
process — on a GitHub Action runner, an out-of-memory crash instead of a score. Inflation
is now capped at 16 MB; real pages never get near it. No score changes.