Skip to content

geo-score 1.8.0

Latest

Choose a tag to compare

@vortwang vortwang released this 05 Oct 09:22
· 7 commits to main since this release

Added

  • How much of the answer your pages support, beside the cited rate. Where a provider documents where each
    citation sits, every answer now records it (spans: start, end, URLs and kind, character offsets into the
    answer as saved, at most 500), read only from documented fields: the start_index/end_index of OpenAI's and
    OpenRouter's url_citation annotations (where the citation sits), Perplexity's [n] marks when its search
    results carry the id its documentation maps them to, Gemini's url_citation annotations (Interactions API)
    and grounding supports (generate_content), both measured in bytes as Google's reference says, and the text
    of Claude's text blocks that carry citations. From them each answer gets supported_share: the share of the
    answer a URL the brand owns supports, by characters, by sentences, and by characters weighted by their
    sentence's position, e^(−position / sentences), after the position-adjusted word count of arXiv 2311.09735,
    §2.2.1, Eq. 3 (characters for words; a sentence several sources support counts in full). A mark supports the
    sentence it closes; a segment, its own characters. The report's new "Answer share your pages support" table
    gives, per engine and next to its cited rate, never merged into it, the answers measured, the median of each
    share and how many are above zero; summary.engines[].supported_share carries the same. An answer with no
    citation, or with one that came without a usable position (an OpenAI-compatible base_url that strips
    annotations looks the same), is not measured, never 0, and so is every Perplexity answer whose results carry
    no id. It is not mention_share: that counts sentences naming the brand, this counts text its pages support.
  • Sources listed and sources used. Perplexity's adapter kept only the sources the text marks with [n] and
    dropped the rest; Gemini's generate_content lists grounding chunks and ties supports to some of them. Such
    answers now record sources_listed and sources_used, and a new table, "Sources listed and used in the
    answer", gives per engine the sources listed and used per answer, used over listed, and the answers that list
    and use a URL of yours (summary.engines[].sources). citations, cited, the rank and the headline rate read
    exactly as before for the same responses, which a test holds against 1.7.1's reading. OpenAI, Claude,
    OpenRouter and Gemini's Interactions API return one list, the citations tied to the text, so they are not
    measured.
  • A run records what it measured in recorded. A run saved before has none, and its report says the answer share
    and the sources were not recorded rather than that the engines gave none; its summary is computed again when it
    is reported, and the file is not rewritten. schema/watch.v1.json registers every new field, closed and capped;
    guide/watch-methodology.md section 9 gives the definitions, which engines give positions and the papers behind
    them; the glossary gains Answer share and Sources listed, sources used. Measurement only: rubric v1.1, every
    tier and every score are unchanged.
  • Repeat consistency per question. When a run asks a question more than once on an engine, the report adds
    "Repeat consistency of cited domains": the Jaccard overlap of the domains cited by every two answers that cite,
    averaged per question, the median per engine, and the least consistent questions (summary.engines[].consistency).
    Answers that cite nothing are left out; one ask per question reads not measured. Descriptive only: watch diff
    does not read it. For scale, the methodology cites arXiv 2604.07585 §5 tables 4 and 6.
  • A precision hint in the plan. watch run --dry-run (and the MCP run tool's dry run) prints, beside the
    resolution, the worst-case standard error of one question's cited rate on one engine at the configured runs
    (0.5/√n), and that more asks add no questions. The methodology adds the external reference point of arXiv
    2604.07585 §7 (at least 7 asks per prompt per day for brands, 8 for sources; one market, one study).
  • Interfaces are never pooled. The methodology (section 7) records that one product's web app, mobile app and
    API cite only partly overlapping sources (arXiv 2607.15771 §4.2, §4.5; no within-interface baseline, §4.1), and
    states the rule: results from different interfaces are never pooled or paired. Every run's channel is api
    today; a manual channel, if added, must record interface, device and sign-in state.
  • A preregistered validation plan (benchmark/PREREGISTRATION.md + zh), written before any data and registered
    by the first tagged release that contains it: does AIV readiness go with being cited (watch's citation rate) and
    with absorption (supported_share.chars, computed for every site of a category from the saved spans; not
    measured where a provider documents no positions)? Sites that compete for the same unbranded questions in 12 to 20
    categories, drawn across the rubric's own band cuts; 8 asks per question per engine per wave (arXiv 2604.07585 §7),
    two waves 12 weeks apart; every site of a category read on the same answers as its concurrent control (arXiv
    2606.04362 §5); every site re-scored and logged each wave (arXiv 2506.11097 §6.4); a within-category permutation
    test and four Holm-corrected secondary analyses written in advance (arXiv 2604.25707 §11.1), and what is published
    whatever the result. It tests the instrument, not whether GEO works; a causal study would randomise page changes
    (arXiv 2604.25707 appendix B). Linked from VALIDITY.md (+ zh) through benchmark/validity.py. No data collected yet.

Changed

  • Evidence table: literature added to five rows, metadata only (rubric/evidence-v1.1.md + zh and the
    evidence objects in rubric/v1.1.json). p2.question-intent adds arXiv 2603.25500 §5.2 Table 6, 2511.20867 §5.3
    Table 5 and 2606.04362 §4: thin, slightly contrary. p2.freshness adds 2607.15771 §2.5 Table 2-5: pages cited for
    highly time-sensitive questions are newer than those cited for low-time-sensitivity ones, which describes the age
    of cited pages, not an effect of showing a date. p2.sourced-stats adds the GEO paper's own all-sources result
    (2311.09735 §5.2 Table 2), the rank-based counter-evidence (2506.11097 §6.2, §6.4) and 2402.11782. p3.listings
    and p3.mentions add 2607.15771 §2.1 Table 2-1 and §3.2 Table 3-2, and state 2606.20065 §6.5 exactly (2.9% own
    domain, 75.2% other companies). Those rows re-verified 2026-10-05. The README's limitations cite the external
    evidence that two asks of one question differ: source overlap of 0.32 to 0.43 per industry with engines pooled and
    0.23 to 0.51 per engine (2604.07585 §5 Tables 4 and 6), 0.66 for AI Overviews (2604.27790 §4.4 Table 4). No weight,
    tier, score or evidence basis changed.

Fixed

  • The README no longer quotes the GEO paper's "up to 40%" without its limits. It now says the gain was
    measured in a simulated engine with one source rewritten at a time, that the top-ranked source lost 20 to 30%
    when every source was rewritten (arXiv 2311.09735 §5.2 Table 2), and that a citation-rank benchmark found most
    methods gave no significant gain (arXiv 2506.11097 §6.2).

Verify downloads: shasum -a 256 -c SHA256SUMS.
中文更新日志:https://github.com/jianruntech/geo-score/blob/v1.8.0/CHANGELOG.zh-CN.md