Repository navigation
Added
- How much of the answer your pages support, beside the cited rate. Where a provider documents where each
citation sits, every answer now records it (spans: start, end, URLs and kind, character offsets into the
answer as saved, at most 500), read only from documented fields: thestart_index/end_indexof OpenAI's and
OpenRouter'surl_citationannotations (where the citation sits), Perplexity's[n]marks when its search
results carry theidits documentation maps them to, Gemini'surl_citationannotations (Interactions API)
and grounding supports (generate_content), both measured in bytes as Google's reference says, and the text
of Claude's text blocks that carry citations. From them each answer getssupported_share: the share of the
answer a URL the brand owns supports, by characters, by sentences, and by characters weighted by their
sentence's position, e^(−position / sentences), after the position-adjusted word count of arXiv 2311.09735,
§2.2.1, Eq. 3 (characters for words; a sentence several sources support counts in full). A mark supports the
sentence it closes; a segment, its own characters. The report's new "Answer share your pages support" table
gives, per engine and next to its cited rate, never merged into it, the answers measured, the median of each
share and how many are above zero;summary.engines[].supported_sharecarries the same. An answer with no
citation, or with one that came without a usable position (an OpenAI-compatiblebase_urlthat strips
annotations looks the same), is not measured, never 0, and so is every Perplexity answer whose results carry
no id. It is notmention_share: that counts sentences naming the brand, this counts text its pages support. - Sources listed and sources used. Perplexity's adapter kept only the sources the text marks with
[n]and
dropped the rest; Gemini'sgenerate_contentlists grounding chunks and ties supports to some of them. Such
answers now recordsources_listedandsources_used, and a new table, "Sources listed and used in the
answer", gives per engine the sources listed and used per answer, used over listed, and the answers that list
and use a URL of yours (summary.engines[].sources).citations,cited, the rank and the headline rate read
exactly as before for the same responses, which a test holds against 1.7.1's reading. OpenAI, Claude,
OpenRouter and Gemini's Interactions API return one list, the citations tied to the text, so they are not
measured. - A run records what it measured in
recorded. A run saved before has none, and its report says the answer share
and the sources were not recorded rather than that the engines gave none; its summary is computed again when it
is reported, and the file is not rewritten. schema/watch.v1.json registers every new field, closed and capped;
guide/watch-methodology.md section 9 gives the definitions, which engines give positions and the papers behind
them; the glossary gains Answer share and Sources listed, sources used. Measurement only: rubric v1.1, every
tier and every score are unchanged. - Repeat consistency per question. When a run asks a question more than once on an engine, the report adds
"Repeat consistency of cited domains": the Jaccard overlap of the domains cited by every two answers that cite,
averaged per question, the median per engine, and the least consistent questions (summary.engines[].consistency).
Answers that cite nothing are left out; one ask per question reads not measured. Descriptive only:watch diff
does not read it. For scale, the methodology cites arXiv 2604.07585 §5 tables 4 and 6. - A precision hint in the plan.
watch run --dry-run(and the MCP run tool's dry run) prints, beside the
resolution, the worst-case standard error of one question's cited rate on one engine at the configuredruns
(0.5/√n), and that more asks add no questions. The methodology adds the external reference point of arXiv
2604.07585 §7 (at least 7 asks per prompt per day for brands, 8 for sources; one market, one study). - Interfaces are never pooled. The methodology (section 7) records that one product's web app, mobile app and
API cite only partly overlapping sources (arXiv 2607.15771 §4.2, §4.5; no within-interface baseline, §4.1), and
states the rule: results from different interfaces are never pooled or paired. Every run'schannelisapi
today; a manual channel, if added, must record interface, device and sign-in state. - A preregistered validation plan (
benchmark/PREREGISTRATION.md+ zh), written before any data and registered
by the first tagged release that contains it: does AIV readiness go with being cited (watch's citation rate) and
with absorption (supported_share.chars, computed for every site of a category from the savedspans; not
measured where a provider documents no positions)? Sites that compete for the same unbranded questions in 12 to 20
categories, drawn across the rubric's own band cuts; 8 asks per question per engine per wave (arXiv 2604.07585 §7),
two waves 12 weeks apart; every site of a category read on the same answers as its concurrent control (arXiv
2606.04362 §5); every site re-scored and logged each wave (arXiv 2506.11097 §6.4); a within-category permutation
test and four Holm-corrected secondary analyses written in advance (arXiv 2604.25707 §11.1), and what is published
whatever the result. It tests the instrument, not whether GEO works; a causal study would randomise page changes
(arXiv 2604.25707 appendix B). Linked from VALIDITY.md (+ zh) throughbenchmark/validity.py. No data collected yet.
Changed
- Evidence table: literature added to five rows, metadata only (
rubric/evidence-v1.1.md+ zh and the
evidenceobjects inrubric/v1.1.json).p2.question-intentadds arXiv 2603.25500 §5.2 Table 6, 2511.20867 §5.3
Table 5 and 2606.04362 §4: thin, slightly contrary.p2.freshnessadds 2607.15771 §2.5 Table 2-5: pages cited for
highly time-sensitive questions are newer than those cited for low-time-sensitivity ones, which describes the age
of cited pages, not an effect of showing a date.p2.sourced-statsadds the GEO paper's own all-sources result
(2311.09735 §5.2 Table 2), the rank-based counter-evidence (2506.11097 §6.2, §6.4) and 2402.11782.p3.listings
andp3.mentionsadd 2607.15771 §2.1 Table 2-1 and §3.2 Table 3-2, and state 2606.20065 §6.5 exactly (2.9% own
domain, 75.2% other companies). Those rows re-verified 2026-10-05. The README's limitations cite the external
evidence that two asks of one question differ: source overlap of 0.32 to 0.43 per industry with engines pooled and
0.23 to 0.51 per engine (2604.07585 §5 Tables 4 and 6), 0.66 for AI Overviews (2604.27790 §4.4 Table 4). No weight,
tier, score or evidence basis changed.
Fixed
- The README no longer quotes the GEO paper's "up to 40%" without its limits. It now says the gain was
measured in a simulated engine with one source rewritten at a time, that the top-ranked source lost 20 to 30%
when every source was rewritten (arXiv 2311.09735 §5.2 Table 2), and that a citation-rank benchmark found most
methods gave no significant gain (arXiv 2506.11097 §6.2).
Verify downloads: shasum -a 256 -c SHA256SUMS.
中文更新日志:https://github.com/jianruntech/geo-score/blob/v1.8.0/CHANGELOG.zh-CN.md