Repository navigation
geo-score 1.7.0
Added
engines[].timeout. Seconds to wait for one engine's answer, 10 to 600 (default 120), for a slow endpoint
or a model that searches and reasons at length. A timed-out request is still not retried.- "What the engines searched for" counts years and ranking words, and names searches that came back without
their queries. A line under the table gives, over every search the answers exposed, how many contain a year
(a four-digit number from 1900 to 2099 standing alone) and how many contain a ranking or comparison word:best,
top,vs,versus,review,reviews,compare,comparison,alternative,alternativesas whole words
in any letter case, or最好,最佳,推荐,排行,排名,哪家,哪个好,对比,评测anywhere. A search
counts once for each, and one that repeats the question's own words counts too: the line describes what was
searched, not what the engine added. When OpenAI, Gemini or Anthropic answers report a search but expose no query,
as an OpenAI-compatiblebase_urlthat returnsweb_search_callitems with noactiondoes, the section no
longer disappears: it says how many answers searched with no query text, per engine, andwatch asksays the same
under the answer. Perplexity and OpenRouter answers are never counted there, since their adapters record one
search per answer whatever happened. The summary gainssummary.searches.answers_searched_without_queries,
summary.searches.modifiers(total,with_year,with_ranking_word) andsearched_without_querieson each
row ofsummary.engines, and schema/watch.v1.json now describessummary.searches. A run saved before this
release gets them computed from its saved answers when it is reported; the run file is not rewritten. - Pages robots.txt blocks by path, and rules two readings disagree on, outside the score.
g.robotsreads only a
Disallow that closes the whole site, so a rule on one path left it green. Each run now holds the homepage, every
sampled page and the first 200 URLs on the audited host and path of the sitemapp1.sitemapread (the first child
of a sitemap index; when the homepage redirects to another host, that host's too, if robots.txt was answered from
there) against the rules each of the ten retrieval crawlers obeys, read as RFC 9309 section 2.2.2 says: the longest
matching pattern decides, an allow wins a tie,*and a final$are wildcards, and percent-encoding is compared
in one form. The report names each page a rule closes, the crawlers and the rule, and the pages where the
first-match reading of Python'surllib.robotparserand many older checkers (rules in file order, the first plain
prefix decides) says otherwise, with the rule each reading chose. Only the first 500 KiB of robots.txt are read, the
least RFC 9309 section 2.5 asks a crawler to parse, and the readings stop after 1,000,000 rule comparisons, counting
the URLs left as not checked, so a file of thousands of wildcard rules cannot hold a run up. Nothing is fetched for
it and it is never scored: rubric v1.1, every tier and the leaderboard are unchanged. - Whether pages are linked in the server HTML, outside the score. From the HTML the run already fetched (the
homepage, the scored sample and the section indexes sampling read), the report counts the<a href>links to pages
on the audited host and path, or on the host the homepage redirected to (apex to www), that a crawler running no
JavaScript can follow: links inside comments, scripts, styles and templates do not count, nor dojavascript:,
mailto:andtel:links. It says how many pages the homepage links to and how many of the other sampled pages
among them, flags a homepage that links to fewer than 3 ("navigation that appears only after JavaScript runs cannot
be followed by crawlers that do not run it"), and names the checked sitemap URLs no fetched page links to (the
homepage aside, since a crawl starts there). Only the fetched pages are searched, so a page they do not link to is
not shown to be orphaned, and the report says so. No request is added, and it is never scored. - hreflang annotations on the fetched pages, outside the score. On the same pages, the report reads each
<link rel="alternate" hreflang="…" href="…">in the server HTML (outside comments, scripts, styles and templates; a
feed's alternate left out) and eachrel="alternate"link with anhreflangparameter in the page's HTTPLink
header, everyLinkline read (a response's repeatedLinklines used to keep only the last). It says how many
pages carry annotations, how many there are and how many came in aLinkheader, and lists, as found: pages that
carry annotations after the end of the head (after</head>or the<body>tag), since Google's hreflang
documentation says the<link>tags must be inside a well-formed<head>(they are still counted with the others);
pages whose annotations leave the page itself out; codes outside the subset of BCP 47 that Google's hreflang
documentation supports (an ISO 639-1 language, optionally an ISO 15924 script and an officially assigned ISO 3166-1
alpha-2 region, orx-default), soen_US,jp,en-UK,es-419and the registry's six deprecated two-letter
language subtags (iwand the like) are listed, each with why; a code that names more than one URL on one page; URLs
named in annotations that are not absolute, as written; and URLs on another host, which hreflang allows and the
report only notes. A>inside a quoted attribute value does not end a<link>, a backslash-escaped quote in a
Linkheader parameter does not end it, and a default port or a trailing dot on a host names the same host. Return
links are checked only between the pages the run fetched: a pair where one page names the other and is not named
back is listed, and the other URLs the annotations name are counted as not checked. A site with none gets one line.
The codes are the 184 two-letter language subtags the IANA Language Subtag Registry (File-Date 2026-09-17) does not
mark deprecated and the 249 officially assigned ISO 3166-1 alpha-2 codes among its region subtags; the registry and
Google's page (last updated 2026-09-21) were read on 2026-10-04. Only the first 1,000 annotations of a page are read, and the rest are counted. No request is added, and
it is never scored: rubric v1.1, every tier and the leaderboard are unchanged. - The text report, the MCP
score_sitesummary and the Action's job summary show all three under "Outside the
score". The JSON report carries them in a new optionalaccessfield (access.robots_paths,
access.server_links,access.hreflang), registered in schema/report.v2.json; each lists the first 50 pages, and
its counts cover them all. The terminal report quotes a page's path, a robots.txt pattern and an hreflang code or
href as«site text: …», the Markdown views as code. A new e2e fixture,hidden-pages, has a path rule that closes
its pricing page, rules the two readings part on, and a menu only a script writes; another,hreflang, has a page
that is not named back, annotations in twoLinkheader lines, and the code and href slips above; thehostile
fixture carries its injection text in an hreflang code and href. Every fixture'sexpected.jsonpins the new
counts. - The Claude Code skill installs with
npx skills add jianruntech/geo-score(the open skills CLI;-gfor a
global install). Checked on 2026-10-04: it finds one skill,geo-score, and copies the whole repository, about
9 MB, into.claude/skills/geo-score. The README'sgit clonestays as the way without Node.js.
Changed
- A ledger with no priced answer says the cost is unknown. It read
$0 + 4 answers with no price (add prices to geo-score-watch.json), which looks like a free run; it now readscost unknown: 4 answers with no price (add prices to geo-score-watch.json). When some answers are priced it still reads$0.07 + 12 answers with no price.
The MCPruntool's message for a cancelled run uses the same words. The JSON ledger is unchanged. - The ledger says what failed. When calls failed,
watch reportadds a line with how many, and how many HTTP
requests were sent in all, since a provider may bill a request it timed out on. A path outside the current
directory is printed absolute instead of as a chain of../.
Fixed
- OpenAI engines send the question as a one-message list. The Responses API accepts a string or a list;
some OpenAI-compatible endpoints that abase_urlpoints at accept only the list and refused every call.
Official OpenAI answers the same either way. - The crawler reference no longer says that blocking an opt-out token costs nothing.
reference/ai-crawlers.md
put the cost of blocking an opt-out token at "nothing, in citation terms" and closed by saying that blocking every
training crawler and both opt-out tokens "costs nothing". Google's common crawlers page (last updated 2026-07-14,
checked 2026-10-04) saysGoogle-Extendedgoverns grounding in Gemini Apps and in Grounding with Google Search on
Vertex AI as well as the training of Gemini models, and does not affect inclusion or ranking in Google Search. The
reference now says what each opt-out token and training crawler controls, citing the vendor's page with its date:
Applebot-Extendedgoverns training only, and Apple's AI answers are opted out of withnosnippet(About Applebot,
published 2026-09-04); Meta's crawler page (undated, checked 2026-10-04) describesMeta-ExternalAgentas crawling
for training and also for indexing content to improve Meta's products, so it is no longer implied to be
training-only;anthropic-aiis not in Anthropic's current crawler article, so whether it sends requests is marked
no data instead of falling under a heading that said opt-out tokens send none. A new table gives the control Google
documents for each part of its AI features, among them Search Console's Search generative AI control for AI
Overviews and AI Mode, and eligibility for those features now follows Google's newer guide (last updated
2026-07-10), which adds inclusion under that control to being indexed and eligible for a snippet. Nothing is scored
differently: training crawlers and opt-out tokens are still not scored,
rubric v1.1 is unchanged, and the glossary's opt-out token entry says the same in both languages. reference/platform-source-selection.mdgave AI Overviews' grounding control toGoogle-Extended, and offered
observations this repository has no data for. Per Google, AI Overviews and AI Mode answer to Googlebot's
robots.txt rules, the snippet controls (nosnippet,data-nosnippet,max-snippet,noindex) and Search
Console's Search generative AI control;Google-Extendedgoverns Gemini's grounding. The page no longer says that a
retrieval crawler disallowed in robots.txt gets nothing: OpenAI and Perplexity say their user-triggered fetchers
(ChatGPT-User,Perplexity-User) may not follow robots.txt. The lines offered as our
observation (pages withOrganizationorArticleschema, recently updated pages, Q&A pages and PDFs cited
disproportionately;llms.txtmattering more to Claude; an unindexed Bing site as the most common blind spot) had
no measurement behind them here, since the benchmark scores sites, not answers, and are gone with the other unsourced
lines; the passage as the unit of selection is now labelled a hypothesis. Each engine's section now cites its
vendor's documentation, re-read on 2026-10-04, and says no data where the vendor documents nothing.
Verify downloads: shasum -a 256 -c SHA256SUMS.
中文更新日志:https://github.com/jianruntech/geo-score/blob/v1.7.0/CHANGELOG.zh-CN.md