Skip to content

Releases: jianruntech/geo-score

geo-score 1.8.0

Choose a tag to compare

@vortwang vortwang released this 05 Oct 09:22

Added

  • How much of the answer your pages support, beside the cited rate. Where a provider documents where each
    citation sits, every answer now records it (spans: start, end, URLs and kind, character offsets into the
    answer as saved, at most 500), read only from documented fields: the start_index/end_index of OpenAI's and
    OpenRouter's url_citation annotations (where the citation sits), Perplexity's [n] marks when its search
    results carry the id its documentation maps them to, Gemini's url_citation annotations (Interactions API)
    and grounding supports (generate_content), both measured in bytes as Google's reference says, and the text
    of Claude's text blocks that carry citations. From them each answer gets supported_share: the share of the
    answer a URL the brand owns supports, by characters, by sentences, and by characters weighted by their
    sentence's position, e^(−position / sentences), after the position-adjusted word count of arXiv 2311.09735,
    §2.2.1, Eq. 3 (characters for words; a sentence several sources support counts in full). A mark supports the
    sentence it closes; a segment, its own characters. The report's new "Answer share your pages support" table
    gives, per engine and next to its cited rate, never merged into it, the answers measured, the median of each
    share and how many are above zero; summary.engines[].supported_share carries the same. An answer with no
    citation, or with one that came without a usable position (an OpenAI-compatible base_url that strips
    annotations looks the same), is not measured, never 0, and so is every Perplexity answer whose results carry
    no id. It is not mention_share: that counts sentences naming the brand, this counts text its pages support.
  • Sources listed and sources used. Perplexity's adapter kept only the sources the text marks with [n] and
    dropped the rest; Gemini's generate_content lists grounding chunks and ties supports to some of them. Such
    answers now record sources_listed and sources_used, and a new table, "Sources listed and used in the
    answer", gives per engine the sources listed and used per answer, used over listed, and the answers that list
    and use a URL of yours (summary.engines[].sources). citations, cited, the rank and the headline rate read
    exactly as before for the same responses, which a test holds against 1.7.1's reading. OpenAI, Claude,
    OpenRouter and Gemini's Interactions API return one list, the citations tied to the text, so they are not
    measured.
  • A run records what it measured in recorded. A run saved before has none, and its report says the answer share
    and the sources were not recorded rather than that the engines gave none; its summary is computed again when it
    is reported, and the file is not rewritten. schema/watch.v1.json registers every new field, closed and capped;
    guide/watch-methodology.md section 9 gives the definitions, which engines give positions and the papers behind
    them; the glossary gains Answer share and Sources listed, sources used. Measurement only: rubric v1.1, every
    tier and every score are unchanged.
  • Repeat consistency per question. When a run asks a question more than once on an engine, the report adds
    "Repeat consistency of cited domains": the Jaccard overlap of the domains cited by every two answers that cite,
    averaged per question, the median per engine, and the least consistent questions (summary.engines[].consistency).
    Answers that cite nothing are left out; one ask per question reads not measured. Descriptive only: watch diff
    does not read it. For scale, the methodology cites arXiv 2604.07585 §5 tables 4 and 6.
  • A precision hint in the plan. watch run --dry-run (and the MCP run tool's dry run) prints, beside the
    resolution, the worst-case standard error of one question's cited rate on one engine at the configured runs
    (0.5/√n), and that more asks add no questions. The methodology adds the external reference point of arXiv
    2604.07585 §7 (at least 7 asks per prompt per day for brands, 8 for sources; one market, one study).
  • Interfaces are never pooled. The methodology (section 7) records that one product's web app, mobile app and
    API cite only partly overlapping sources (arXiv 2607.15771 §4.2, §4.5; no within-interface baseline, §4.1), and
    states the rule: results from different interfaces are never pooled or paired. Every run's channel is api
    today; a manual channel, if added, must record interface, device and sign-in state.
  • A preregistered validation plan (benchmark/PREREGISTRATION.md + zh), written before any data and registered
    by the first tagged release that contains it: does AIV readiness go with being cited (watch's citation rate) and
    with absorption (supported_share.chars, computed for every site of a category from the saved spans; not
    measured where a provider documents no positions)? Sites that compete for the same unbranded questions in 12 to 20
    categories, drawn across the rubric's own band cuts; 8 asks per question per engine per wave (arXiv 2604.07585 §7),
    two waves 12 weeks apart; every site of a category read on the same answers as its concurrent control (arXiv
    2606.04362 §5); every site re-scored and logged each wave (arXiv 2506.11097 §6.4); a within-category permutation
    test and four Holm-corrected secondary analyses written in advance (arXiv 2604.25707 §11.1), and what is published
    whatever the result. It tests the instrument, not whether GEO works; a causal study would randomise page changes
    (arXiv 2604.25707 appendix B). Linked from VALIDITY.md (+ zh) through benchmark/validity.py. No data collected yet.

Changed

  • Evidence table: literature added to five rows, metadata only (rubric/evidence-v1.1.md + zh and the
    evidence objects in rubric/v1.1.json). p2.question-intent adds arXiv 2603.25500 §5.2 Table 6, 2511.20867 §5.3
    Table 5 and 2606.04362 §4: thin, slightly contrary. p2.freshness adds 2607.15771 §2.5 Table 2-5: pages cited for
    highly time-sensitive questions are newer than those cited for low-time-sensitivity ones, which describes the age
    of cited pages, not an effect of showing a date. p2.sourced-stats adds the GEO paper's own all-sources result
    (2311.09735 §5.2 Table 2), the rank-based counter-evidence (2506.11097 §6.2, §6.4) and 2402.11782. p3.listings
    and p3.mentions add 2607.15771 §2.1 Table 2-1 and §3.2 Table 3-2, and state 2606.20065 §6.5 exactly (2.9% own
    domain, 75.2% other companies). Those rows re-verified 2026-10-05. The README's limitations cite the external
    evidence that two asks of one question differ: source overlap of 0.32 to 0.43 per industry with engines pooled and
    0.23 to 0.51 per engine (2604.07585 §5 Tables 4 and 6), 0.66 for AI Overviews (2604.27790 §4.4 Table 4). No weight,
    tier, score or evidence basis changed.

Fixed

  • The README no longer quotes the GEO paper's "up to 40%" without its limits. It now says the gain was
    measured in a simulated engine with one source rewritten at a time, that the top-ranked source lost 20 to 30%
    when every source was rewritten (arXiv 2311.09735 §5.2 Table 2), and that a citation-rank benchmark found most
    methods gave no significant gain (arXiv 2506.11097 §6.2).

Verify downloads: shasum -a 256 -c SHA256SUMS.
中文更新日志:https://github.com/jianruntech/geo-score/blob/v1.8.0/CHANGELOG.zh-CN.md

geo-score 1.7.1

Choose a tag to compare

@vortwang vortwang released this 04 Oct 15:44

Fixed

  • A run started at one page no longer reads its site navigation as no links. The server-HTML links
    observation (new in 1.7.0) kept only links inside the path the run was given, so geo_score.py example.com/product
    counted 0 links on a page with a full navigation bar and flagged it as navigation that appears only after
    JavaScript runs. Links now count anywhere on the host, the start page itself aside, and the report names that page
    as the page the run started from rather than the homepage ("links to N pages on this host"). Reported from a live
    site the day 1.7.0 shipped. Never part of the score.

Verify downloads: shasum -a 256 -c SHA256SUMS.
中文更新日志:https://github.com/jianruntech/geo-score/blob/v1.7.1/CHANGELOG.zh-CN.md

geo-score 1.7.0

Choose a tag to compare

@vortwang vortwang released this 04 Oct 15:03

Added

  • engines[].timeout. Seconds to wait for one engine's answer, 10 to 600 (default 120), for a slow endpoint
    or a model that searches and reasons at length. A timed-out request is still not retried.
  • "What the engines searched for" counts years and ranking words, and names searches that came back without
    their queries.
    A line under the table gives, over every search the answers exposed, how many contain a year
    (a four-digit number from 1900 to 2099 standing alone) and how many contain a ranking or comparison word: best,
    top, vs, versus, review, reviews, compare, comparison, alternative, alternatives as whole words
    in any letter case, or 最好, 最佳, 推荐, 排行, 排名, 哪家, 哪个好, 对比, 评测 anywhere. A search
    counts once for each, and one that repeats the question's own words counts too: the line describes what was
    searched, not what the engine added. When OpenAI, Gemini or Anthropic answers report a search but expose no query,
    as an OpenAI-compatible base_url that returns web_search_call items with no action does, the section no
    longer disappears: it says how many answers searched with no query text, per engine, and watch ask says the same
    under the answer. Perplexity and OpenRouter answers are never counted there, since their adapters record one
    search per answer whatever happened. The summary gains summary.searches.answers_searched_without_queries,
    summary.searches.modifiers (total, with_year, with_ranking_word) and searched_without_queries on each
    row of summary.engines, and schema/watch.v1.json now describes summary.searches. A run saved before this
    release gets them computed from its saved answers when it is reported; the run file is not rewritten.
  • Pages robots.txt blocks by path, and rules two readings disagree on, outside the score. g.robots reads only a
    Disallow that closes the whole site, so a rule on one path left it green. Each run now holds the homepage, every
    sampled page and the first 200 URLs on the audited host and path of the sitemap p1.sitemap read (the first child
    of a sitemap index; when the homepage redirects to another host, that host's too, if robots.txt was answered from
    there) against the rules each of the ten retrieval crawlers obeys, read as RFC 9309 section 2.2.2 says: the longest
    matching pattern decides, an allow wins a tie, * and a final $ are wildcards, and percent-encoding is compared
    in one form. The report names each page a rule closes, the crawlers and the rule, and the pages where the
    first-match reading of Python's urllib.robotparser and many older checkers (rules in file order, the first plain
    prefix decides) says otherwise, with the rule each reading chose. Only the first 500 KiB of robots.txt are read, the
    least RFC 9309 section 2.5 asks a crawler to parse, and the readings stop after 1,000,000 rule comparisons, counting
    the URLs left as not checked, so a file of thousands of wildcard rules cannot hold a run up. Nothing is fetched for
    it and it is never scored: rubric v1.1, every tier and the leaderboard are unchanged.
  • Whether pages are linked in the server HTML, outside the score. From the HTML the run already fetched (the
    homepage, the scored sample and the section indexes sampling read), the report counts the <a href> links to pages
    on the audited host and path, or on the host the homepage redirected to (apex to www), that a crawler running no
    JavaScript can follow: links inside comments, scripts, styles and templates do not count, nor do javascript:,
    mailto: and tel: links. It says how many pages the homepage links to and how many of the other sampled pages
    among them, flags a homepage that links to fewer than 3 ("navigation that appears only after JavaScript runs cannot
    be followed by crawlers that do not run it"), and names the checked sitemap URLs no fetched page links to (the
    homepage aside, since a crawl starts there). Only the fetched pages are searched, so a page they do not link to is
    not shown to be orphaned, and the report says so. No request is added, and it is never scored.
  • hreflang annotations on the fetched pages, outside the score. On the same pages, the report reads each
    <link rel="alternate" hreflang="…" href="…"> in the server HTML (outside comments, scripts, styles and templates; a
    feed's alternate left out) and each rel="alternate" link with an hreflang parameter in the page's HTTP Link
    header, every Link line read (a response's repeated Link lines used to keep only the last). It says how many
    pages carry annotations, how many there are and how many came in a Link header, and lists, as found: pages that
    carry annotations after the end of the head (after </head> or the <body> tag), since Google's hreflang
    documentation says the <link> tags must be inside a well-formed <head> (they are still counted with the others);
    pages whose annotations leave the page itself out; codes outside the subset of BCP 47 that Google's hreflang
    documentation supports (an ISO 639-1 language, optionally an ISO 15924 script and an officially assigned ISO 3166-1
    alpha-2 region, or x-default), so en_US, jp, en-UK, es-419 and the registry's six deprecated two-letter
    language subtags (iw and the like) are listed, each with why; a code that names more than one URL on one page; URLs
    named in annotations that are not absolute, as written; and URLs on another host, which hreflang allows and the
    report only notes. A > inside a quoted attribute value does not end a <link>, a backslash-escaped quote in a
    Link header parameter does not end it, and a default port or a trailing dot on a host names the same host. Return
    links are checked only between the pages the run fetched: a pair where one page names the other and is not named
    back is listed, and the other URLs the annotations name are counted as not checked. A site with none gets one line.
    The codes are the 184 two-letter language subtags the IANA Language Subtag Registry (File-Date 2026-09-17) does not
    mark deprecated and the 249 officially assigned ISO 3166-1 alpha-2 codes among its region subtags; the registry and
    Google's page (last updated 2026-09-21) were read on 2026-10-04. Only the first 1,000 annotations of a page are read, and the rest are counted. No request is added, and
    it is never scored: rubric v1.1, every tier and the leaderboard are unchanged.
  • The text report, the MCP score_site summary and the Action's job summary show all three under "Outside the
    score".
    The JSON report carries them in a new optional access field (access.robots_paths,
    access.server_links, access.hreflang), registered in schema/report.v2.json; each lists the first 50 pages, and
    its counts cover them all. The terminal report quotes a page's path, a robots.txt pattern and an hreflang code or
    href as «site text: …», the Markdown views as code. A new e2e fixture, hidden-pages, has a path rule that closes
    its pricing page, rules the two readings part on, and a menu only a script writes; another, hreflang, has a page
    that is not named back, annotations in two Link header lines, and the code and href slips above; the hostile
    fixture carries its injection text in an hreflang code and href. Every fixture's expected.json pins the new
    counts.
  • The Claude Code skill installs with npx skills add jianruntech/geo-score (the open skills CLI; -g for a
    global install). Checked on 2026-10-04: it finds one skill, geo-score, and copies the whole repository, about
    9 MB, into .claude/skills/geo-score. The README's git clone stays as the way without Node.js.

Changed

  • A ledger with no priced answer says the cost is unknown. It read $0 + 4 answers with no price (add prices to geo-score-watch.json), which looks like a free run; it now reads cost unknown: 4 answers with no price (add prices to geo-score-watch.json). When some answers are priced it still reads $0.07 + 12 answers with no price.
    The MCP run tool's message for a cancelled run uses the same words. The JSON ledger is unchanged.
  • The ledger says what failed. When calls failed, watch report adds a line with how many, and how many HTTP
    requests were sent in all, since a provider may bill a request it timed out on. A path outside the current
    directory is printed absolute instead of as a chain of ../.

Fixed

  • OpenAI engines send the question as a one-message list. The Responses API accepts a string or a list;
    some OpenAI-compatible endpoints that a base_url points at accept only the list and refused every call.
    Official OpenAI answers the same either way.
  • The crawler reference no longer says that blocking an opt-out token costs nothing. reference/ai-crawlers.md
    put the cost of blocking an opt-out token at "nothing, in citation terms" and closed by saying that blocking every
    training crawler and both opt-out tokens "costs nothing". Google's common crawlers page (last updated 2026-07-14,
    checked 2026-10-04) says Google-Extended governs grounding in Gemini Apps and in Grounding with Google Search on
    Vertex AI as well as the training of Gemini models, and does not affect inclusion or ranking in Google Search. The
    reference now says what each opt-out token and training crawler controls, citing the vendor's page with its date:
    Applebot-Extended governs training only, and Apple's AI answers are opted out of with nosnippet (About Applebot,
    published 2026-09-04); Meta's crawler page (undated, checked 2026-10-04) describes Meta-ExternalAgent as crawling
    for training and also for indexing content to improve Meta's products, so it is no longer implied to be
    training-only; anthropic-ai is not in Anthropic's current crawler article, so whether it sends requests is marked
    no data instead of falling under a heading that said opt-out tokens se...
Read more

geo-score 1.6.0

Choose a tag to compare

@vortwang vortwang released this 02 Oct 19:15

Added

  • watch report shows who is named for each question. A new table, "Who is named for each question",
    lists every question with its answers and, out of those, how many cite and name you, how many cite and name
    each competitor, how many name no tracked brand, how many searched the web (yes, no or unknown), and the
    tracked brand cited most often, given only when the question has at least 4 answers and one brand is ahead of
    every other. Every value is a count. Whether an answer searched is derived from what every run already saves
    (a reported search, an exposed search query or a citation); Perplexity and OpenRouter answers are always
    unknown, because their adapters record one search per answer whatever happened. Three lines under the table
    say what the counts cannot: an answer given without a search cannot cite anything, "no tracked brand named"
    covers only the brands in your config, and why some answers are unknown. The summary gains
    summary.by_question ($defs/by_question in schema/watch.v1.json), the JSON report and the MCP report tool
    carry it, and it is an optional property of the tool's output schema, so the MCP contract only grows. A run
    saved before this release gets it computed from its saved answers when it is reported; the run file is not
    rewritten.
  • Suggested question types, and watch check counts them. intent_tier stays free text; cli/README.md and
    guide/watch-methodology.md now suggest eight labels, each with one line and an example: list, vs,
    alternative, pricing, problem (a situation that names no category or brand), persona (who it is for),
    verdict (is it worth it, can it be trusted) and channel (franchise, wholesale, distributor, procurement).
    watch check prints how many questions have each type, with a note, never a failure, when every question has
    the same type or none has one. watch init writes two more placeholder questions, a problem row and a
    verdict row.
  • target_url is read: watch report counts how often the page a question tracks is cited. The column was
    documented and saved but nothing used it. It is now resolved when the questions load: an absolute http(s) URL
    as written (fragment and utm_* parameters dropped), or a path starting with / joined to the brand's first
    domain. watch check, watch run and the MCP run (its dry run too) and status tools warn about any other
    value, which the run ignores, and about a page that is not on the brand's domains or URL prefixes, which is still
    matched. Each answer records target_cited (one of its
    citations is that page) and other_own_cited (it cites another URL of the brand's, not that page), both null
    when the question tracks no page. The match is the exact page once www, letter case, a trailing slash, the
    query string, the fragment and percent-encoding are set aside; /en/x is not /x, and a citation left as an
    unresolved redirect counts as neither. A new table, "Pages you track", gives per page its questions, their
    answers, how many cite the page, how many cite another of your URLs instead, and the URL of yours those
    answers cite most, exactly as cited. The CSV export gains a last column, target_cited (Y, N or empty).
    Schema: target_cited and other_own_cited in $defs/sample, summary.tracked_pages with
    $defs/tracked_page, and a description of the saved target_url. The MCP report tool's output schema gains
    the optional tracked_pages, so the MCP contract only grows. A run saved before this release gets the table
    computed from its saved citations when it is reported.
  • watch report shows who is named first. Each answer records, for the brand and each competitor,
    mention_order (1 for the brand named before every other brand in the config, by the earliest match of its name,
    aliases or domains with the matcher of mentioned; null when it is not named, or when two brands are first named
    at the same place) and mention_share (the share of the answer's sentences that name it, every sentence counted
    the same; a sentence ends at a line break, at 。!?, or at . ! ? before whitespace). With competitors in the
    config, "By engine" gains a column, "named first (tracked)": the answers that name you first among the brands in
    the config, out of the answers that name at least one of them. "Share of voice" gains "named first (tracked)" and
    "sentence share", the mean mention_share over the answers that name each brand, with three lines on how they are
    read. Brands the config does not list are not seen, so first among the brands you track may not be first in the
    answer. Both are read from the answer as saved, which keeps its first 20,000 characters. Schema: mention_order
    and mention_share in $defs/sample and its competitors, and named_first, any_named and mention_share on
    the rows of summary.engines and summary.share_of_voice. The MCP tools' schemas are unchanged. A run saved
    before this release gets them computed from its saved answers when it is reported.
  • watch run --dry-run says what the plan can detect, and watch report can pool runs. The dry run prints the
    plan's resolution: with E engines watch diff tests E + 1 rows (one row with one engine), and it can call a
    change only when at least k questions move the same way between two runs, k being the smallest n with rows × 2 /
    2^n < 0.05, the screen diff already applies: 6 questions with one engine, 7 with two, 8 with three to five. A set
    with fewer questions gets "this plan cannot detect a change": more questions help, more engines raise k, and more
    asks per question do not lower it. The MCP run tool's dry run prints the same lines. watch report --since DAYS
    (or --runs A..B) counts the saved runs of a period as one set of answers, when each asked the same questions
    (the same questions_hash) as many times each (the same runs, so every run weighs the same) on the same engines
    with the same models, from the same country, with the same brand domains and URL prefixes as the newest;
    otherwise it exits 1 and names every run that differs and why. It gives the runs, the calendar days they cover,
    the answers, the headline rate and a row per engine, with a 95% interval that resamples questions, each with every
    answer it got in every run, and the run report's notes (the API channel, questions that name the brand, what
    counts as not measured, and that it does not predict traffic, rankings or revenue). There is no pooled diff.
    Schema: $defs/pooled in schema/watch.v1.json. MCP: the report tool gains an optional since argument and its
    output schema the optional runs and days, so the MCP contract only grows. guide/watch-methodology.md gains
    section 8, "One ask, one day, one month".

Changed

  • Install from PyPI. geo-score has been on PyPI since 1.5.0, as the same wheel and sdist the GitHub release
    attaches, and the install lines now lead with it: uvx geo-score@X.Y.Z example.com,
    pipx install geo-score==X.Y.Z, and uvx geo-score@X.Y.Z mcp for MCP clients (the install buttons, the Claude
    Code line, the standard config, every client in guide/mcp.md and the Claude Desktop bundle's watch-init hint). The
    release wheel URL stays as the line for a machine without PyPI, the git tag as the from-source line, and the curl
    line as the way to run level 1 with nothing installed. server.json lists the pypi package beside the .mcpb, so
    the MCP Registry offers it from this release on, and README.md, the package's description on PyPI, carries the
    registry's mcp-name line (1.5.0's description on PyPI has none, so a 1.5.0 listing names only the .mcpb).
    release.py --bump moves every geo-score@X.Y.Z and geo-score==X.Y.Z, and --execute ends with the PyPI
    upload (--no-pypi leaves it out, --publish-pypi does it alone): it uploads exactly the wheel and sdist that
    SHA256SUMS names, with twine pinned in TWINE_VERSION, then checks PyPI's SHA-256 of each (RELEASING.md,
    step 8). Its requests to GitHub and PyPI are tried three times, and a stop that a rerun can get past names the
    command.
    The message watch, mcp and --ask print when geo_watch.py is missing (the piped one-line run) names the
    package first too, then the wheel.

Fixed

  • watch diff compares a question only with itself. It paired questions by query_id alone, so a question
    reworded under the same id was compared with its old wording, and the edit could read as a change (and make
    --fail-on-drop exit 4). The question text is now compared case-folded, with whitespace collapsed: a question
    whose text changed is left out of every row and named in the notes (at most 10 ids, then "and N more"). The
    6-question minimum and the screen for rows no test could call apply to the questions left, so on runs with
    edited questions a verdict can change, and can become too few questions to tell. The notes also say when the
    runs asked from different locations (country) and when an engine id moved to another provider. Run files gain
    an optional top-level questions_hash (SHA-256 over the sorted pairs of id and normalised text, defined in
    schema/watch.v1.json), and watch runs shows its first 8 characters; for older runs it is computed from the
    saved questions.
  • An API that changes its answer format no longer reads, without a word, as an engine that stopped citing.
    Each adapter reads one format per provider, so a renamed citation field gave answers with no citations at all,
    and watch diff --fail-on-drop could exit 4 on what was a format change. watch report now warns under "By
    engine" when an engine answered 5 or more questions and none of its answers carried a citation of any domain.
    watch diff adds a warning to an engine's row, the same words to the notes and to stderr, when the share of
    its answers carrying a...
Read more

geo-score 1.5.0

Choose a tag to compare

@vortwang vortwang released this 02 Oct 16:39

A release about reading sites the way the rubric says, and about failing CI for the right reason. The two
crawler gates now read only the ten retrieval crawlers, choose each crawler's robots.txt group as RFC 9309
does and retry a request that got no answer before they count it; a bot-challenge page is no longer scored
as a page that needs JavaScript; every page is read in its own language; and a figure counts as sourced
only when its source sits beside it. These fixes move scores, and each one below names its check and its
direction. For CI there are --fail-on-gate and --assert (a gate at zero caps a score at 40, and a site
that would otherwise score 40 or more reads exactly 40, which --fail-under 40 lets through),
geo-score diff and --baseline to compare two reports check by check, and a shields.io endpoint badge.
A site that cannot be scored fails fast and says why in one line, in JSON under --json, and a scoring
run can no longer hang. Around the CLI: a schema and typo warnings for the watch config, installs from the
release wheel without git, a gate that pins its tools, CI examples beyond GitHub, forms for disputing a
leaderboard row, and a benchmark that can be re-measured and taken apart. The leaderboard itself is still
the 1.3.0 run of 2026-09-27.

Upgrading from 1.4.0. Level-1 scores move only where 1.4.0 read a site wrong, and every such fix is
under Fixed with its check and direction: the gates (g.robots, g.reachable, and g.ssr for sites
behind a bot challenge, with Japanese or Korean pages, or with English pages that carry more than 20 Han
characters), the checks that read page text (p2.answer-passages, p2.question-intent,
p2.sourced-stats, p2.named-author, p4.answer-shape, p4.cn-engines and p3.knowledge-graph), and
the checks whose file, logo or crawler request is now retried (p1.llms-txt, b.llms-full, b.ai-txt,
p1.organization and p4.cn-engines). Three invocations exit differently: a site that serves only
bot-challenge pages exits 2 instead of being scored; an unexpected error exits 2 in one line instead of
a traceback that exited 1; and --sample below 1 exits 2, where 1.4.0 crashed on 0 and scored fewer
pages on a negative value. diff is now a subcommand, so a host named diff is written
https://diff. audience_language can hold languages other than en and zh-CN, and the report gains
only the optional elapsed_s.

Added

  • A glossary and a Chinese CLI reference. guide/glossary.md (and guide/glossary.zh-CN.md) defines the 49
    terms the README, the reports and the guides use, each with where it appears and which page goes deeper;
    cli/README.zh-CN.md is the CLI reference in Simplified Chinese, with a language switch at the top of both.
  • --fail-on-gate and --assert EXPR, in the CLI and as the Action inputs fail-on-gate and
    assert. A gate at zero caps the score at 40: a site that would otherwise score 40 or more reads
    exactly 40 and passes --fail-under 40 (fail-under: 40). --fail-on-gate fails any capped site: it
    exits 1 when any g.* check scores 0, and names the gate and its evidence; it is the same as
    --assert 'g.*>0'. --assert holds one check, or an fnmatch glob of checks, to a value, e.g.
    g.robots=max or p1.llms-txt>=4, and is repeatable. A check the run could not observe is skipped,
    never failed. A malformed expression, or a glob that names no check, exits 2.
    Both are opt-in and change no score, and their results go to stderr, never into the report.
  • geo-score diff BEFORE AFTER [--format text|md|json] [--fail-on-drop] compares two level-1 reports
    check by check. Each check shows its points before and after, its max, and a state: up, down,
    same, became unobservable, became observable, or unobservable when neither report observed it.
    The comparison also shows the score, band and gate cap before and after, names each gate that reached
    or left 0, and lists the pages added to or removed from the sample. When the two samples differ, it
    says so, reads the totals as ±5 (benchmark/REPRODUCIBILITY.md) and names g.reachable as the least
    stable check. A check that became unobservable is never shown as a drop to 0. --format json follows
    the new schema/diff.v1.json. Two reports of different rubric versions exit 2, and reports from
    different releases, or of two different targets, are compared with a warning on stderr. diff needs
    only geo_score.py.
  • --baseline REPORT on a scoring run scores the baseline's pages again (as --urls-from does) and
    prints the same comparison after the report. It cannot be combined with --urls or --urls-from.
    Under --json, stdout carries only the report and the comparison is not printed: write the report
    with --json-out and run geo-score diff on the two files to see it. --fail-on-drop still exits 4
    and names each drop on stderr.
  • --fail-on-drop (with --baseline, or on diff) exits 4 when a check lost points, a gate reached
    0 or the band fell, and prints one stderr line per reason. It is opt-in, as watch diff --fail-on-drop
    is, so no existing invocation changes its exit code. A failed --fail-under, --fail-on-gate or
    --assert still exits 1.
  • Action: a baseline input, delta and dropped-checks outputs, and a "Changes since baseline"
    section in the job summary, with the check ids read from the baseline escaped. The baseline may be the
    same file as json-out.
  • A badge that stays current. --badge-json FILE writes the badge as shields.io
    endpoint
    JSON:
    {"schemaVersion": 1, "label": "AIV readiness", "message": "71/100 Solid", "color": "#2f8f52", "cacheSeconds": 86400}.
    The SVG and the JSON take their message and colour from one helper, so the two cannot disagree. Host
    the file anywhere and https://img.shields.io/endpoint?url=<its URL> draws the badge. The Action gets
    a badge input: a path ending in .json gets the endpoint JSON, and any other path gets the SVG. A
    badge-path output says where the file went.
  • A JSON error instead of nothing. Under --json, every exit 2, usage errors included, prints
    {"error": {"kind", "message", "target"}, "tool": "geo-score-cli/X.Y.Z"} on stdout, where stdout used
    to be empty; diff --format json does the same. The new schema/error.v1.json names the kinds: dns,
    refused, tls, timeout, dropped and no_response when nothing answered; http_<code>,
    no_content, slow and bot_challenge when the site answered; deadline when the run ran out of
    time; not_public under the MCP server's address guard; and usage, rubric_mismatch (reports of two
    rubric versions given to diff or --baseline), error and internal. A minor release may add
    kinds; read an unknown one as error.
  • --timeout SECONDS (default 15, as before; up to 300): how long each request may wait for the
    server before it counts as unanswered.
  • elapsed_s in the JSON report (optional, schema/report.v2.json): how many seconds the run took
    on the machine that ran it. The footer prints it too (· took 23.4 s). It describes the run, not the
    site, and diff ignores it.
  • A schema for geo-score-watch.json, and a warning for every key nothing reads.
    schema/watch-config.v1.json (JSON Schema draft 2020-12) describes the file, and watch init now
    writes "$schema" as its first key, so an editor checks and completes it. A mistyped key used to
    vanish without a word: budget for budget_usd meant no budget, and max_call for max_calls the
    default cap of 120. watch check names every unknown key at every level (top, brand,
    competitors[i], engines[i], prices.<id>), with the nearest known key when one is close, as in
    unknown key 'budget' (did you mean 'budget_usd'?); it is ignored, and flags a prices entry that
    names no engine. watch run and the other commands that read the config print the same warnings on
    stderr; the MCP status tool lists them, and run shows them with its plan and after a saved run. A
    key named like a credential (api_key, OPENAI_KEY, api_token, client_secret, credentials …)
    is pointed at environment variables instead, and its value is never printed. Warnings only: no exit
    code changes.
  • CI beyond GitHub. examples/ci/gitlab-ci.yml is a GitLab job on python:3.12-slim: it downloads
    geo_score.py from a pinned release, checks it against SHA256SUMS, scores once, keeps
    geo-score.json as an artifact and hands geo-score.xml to GitLab as a JUnit report, so a check at
    zero shows as a failed test. examples/ci/README.md gives the one directive that Bitbucket Pipelines,
    Azure Pipelines, Jenkins and CircleCI each need to show the --junit file.
  • Where to report what. A Benchmark correction form disputes a leaderboard row's score, gate
    reason, label or not-scored reason, with a fresh --json-out report and the tool version; a Tool
    problem
    form covers the CLI, the Action, the MCP server, pip/pipx/uvx, the MCPB bundle, the Claude
    Code plugin, the Gemini CLI extension and Docker. The issue chooser links the troubleshooting guide and
    the MCP troubleshooting section, and a new SUPPORT.md says which form takes what. The forms applied
    benchmark, engine and rubric, labels the repository did not have, so those issues arrived
    unlabelled: .github/labels.json now lists every label, and scripts/sync_labels.py carries the list
    to GitHub (a dry run by default; --apply creates and updates, --check exits 1 when a label is
    missing, and it never deletes one). It prints a proxy's host and port, never the credentials in
    HTTPS_PROXY.
  • Benchmark: a re-runnable retest study. benchmark/retest.py takes two or more run directories
    (results.json plus raw/) and records each run's tool, date, Python, declared vantage and pinning....
Read more

geo-score 1.4.0

Choose a tag to compare

@vortwang vortwang released this 27 Sep 18:56

The MCP server becomes a first-class way to use geo-score: its own README section with install
buttons, a guide for every client, one-click packages, both protocol eras, and a surface that says
what costs money and lets you turn that off.

Added

  • MCP server, both protocol eras. 2026-07-28 (server/discover, per-request _meta, cache hints)
    alongside 2025-11-25, 2025-06-18, 2025-03-26 and 2024-11-05 over initialize. An unknown legacy
    version gets 2025-11-25; an unsupported modern one gets -32022 with the supported list. Batches are
    accepted only on 2025-03-26, and a cursor is an error because the server never issues one.
  • MCP tools say what they cost. score_site is listed first. Every tool has a title and all four
    annotation hints, descriptions open with the cost ("Free, no API key." / "Paid: …"), input schemas
    reject unknown keys, and nullable fields use anyOf, which strict clients accept. ask and run
    now return structured results, and a run dry run returns its plan as data.
  • Cancelling stops spending. notifications/cancelled stops ask and run before their next
    call; a cancelled run is not saved, and its calls and cost go to stderr. Tool calls run on worker
    threads (up to 4), so the server still answers ping during a long run. At end of input it stops the
    paid calls at once and lets free ones answer. One run at a time: a second is refused while one is in
    progress, so parallel calls cannot double the spend of one approval.
  • Robustness. A malformed run file or config is an error answer, never the end of the server;
    arguments that are not an object are refused rather than read as {}; stdio is UTF-8 on every
    platform, and a stream that cannot carry a character gets it escaped instead of losing the answer;
    status never shows a value pasted into key_env and reports run as ready only when it could start;
    GEO_SCORE_MCP_EXCLUDE_TOOLS and --exclude-tools add up.
  • Choose the tools. --toolsets score,ask,watch, --tools, --exclude-tools and --read-only,
    each with a GEO_SCORE_MCP_* environment twin. --read-only gives a server that cannot spend. A new
    read-only status tool reports the version, the config it found, which engines have keys (never the
    keys) and the caps.
  • Prompts and completions. audit_site, check_citations, weekly_watch, compare_runs and
    explain_check, with completion for check ids and run ids, and a geo-score://runs/{run_id} resource.
    None of them contains fix instructions: measurement is public, remediation is not.
  • Packages. Every release attaches a Claude Desktop bundle (geo-score-X.Y.Z.mcpb) and a
    server.json for the MCP Registry, both covered by SHA256SUMS. The repository is also a Claude Code
    plugin marketplace (/plugin marketplace add jianruntech/geo-score) and a Gemini CLI extension, and
    has a Dockerfile for the stdio server. Nothing is published to a registry, PyPI or a container
    registry yet.
  • Docs. A ## MCP server section in both READMEs with install buttons (Cursor, VS Code, VS Code
    Insiders, LM Studio), a standard config, the tool table and example prompts; guide/mcp.md with the
    exact setup for Claude Code, Claude Desktop, Codex, Cursor, VS Code, Gemini CLI, Devin Desktop and
    Zed, plus timeouts, security and troubleshooting. scripts/mcp_docs.py generates the tables and
    buttons from the server, and check.sh fails when they are stale.
  • Checks. scripts/check_mcp.py drives the server over raw stdio in both eras on every tested
    Python; check.sh --live repeats it through the official MCP Inspector (2.8.0, --strict) and calls
    score_site. scripts/test_packaging.py checks every manifest against the code. A snapshot and a
    1.3.0 contract baseline stop a tool or a required argument from changing by accident.
  • SECURITY.md has a per-tool table of network access, files and spending for the MCP server.

Changed

  • The leaderboard is re-run with 1.3.0 (2026-09-27): 317 of 370 sites scored, median 56 (95% CI
    52–58), range 11–99; 83 sites (26%) have a gate at zero (47 need JavaScript, 23 block AI crawlers,
    13 serve an error). Across the 312 sites scored in both runs the mean change is −0.4 points; 72% move
    by 5 points or less, which mixes the 1.3.0 scoring fixes with 18 days of site changes and is not a
    reproducibility figure (that study is benchmark/REPRODUCIBILITY.md, same tool, two runs).
  • Every site's full report from the run is in benchmark/raw/, and validate.py checks it against
    results.json (rule 29).
  • The benchmark README's findings are computed rather than written by hand: the headline share,
    the top gaps, and which checks drive the Chinese-market gap. That gap is 20 points (95% CI 16–28),
    and its two largest per-check drivers are heuristics the CLI also read lower than a human on the
    Chinese site in the hand-audit comparison; the README says so.
  • The READMEs link the leaderboard on www.jianruntech.com/leaderboard, which is kept current; the
    GitHub Pages copy cannot be rebuilt while Actions and Pages are disabled for the organisation.

Verify downloads: shasum -a 256 -c SHA256SUMS.
中文更新日志:https://github.com/jianruntech/geo-score/blob/v1.4.0/CHANGELOG.zh-CN.md

geo-score 1.3.0

Choose a tag to compare

@vortwang vortwang released this 27 Sep 13:25

A release about trust in the measurement: outputs CI systems read natively, a pinned sample for
before/after comparisons, change detection that no longer calls noise a change, offline end-to-end
tests on twelve fixture sites, a measured CLI-vs-hand-audit agreement, and an evidence basis for every
check. The fixtures also found places where the CLI read sites against the rubric (grouped
robots.txt rules, single-page apps' catch-all pages, empty markup, stale dateModified, bylines,
Chinese dates and figures); those are fixed and listed below with their direction. The percentage can
no longer pass 100, and the MCP server now keeps every connection on public addresses.

Upgrading from 1.2.0. Level 1 scores move only where a check read a site wrongly; each fix and
its direction is listed under Fixed, and the public leaderboard (measured with 1.1.0) is not
re-scored. In watch, the headline rate leaves out questions that name the brand, intervals widen
when a question has several answers, and some diff verdicts that said change now say
within noise or too few questions to tell; the 1.2.0 fields keep their meaning. Evidence strings
now quote site text inside «site text: …». The CSV report has an eleventh column, branded.

Added

  • One run, several outputs: --json-out FILE, --sarif FILE (SARIF 2.1.0: one rule per check,
    one result per check below full marks) and --junit FILE (one test per check; a check at zero
    fails, an unobservable one is skipped). All are written from the same fetch as the report on
    screen, so they cannot disagree with it. Single site only.
  • Install as a command: pyproject.toml, so pipx install git+https://github.com/jianruntech/geo-score
    and uvx --from git+… geo-score work. Installs geo-score and geo-score-mcp (the MCP server over
    stdio). Still no dependencies.
  • STABILITY.md: what is a public contract (check ids, score arithmetic, the JSON schemas, CLI
    flags, exit codes, Action inputs and outputs, MCP tool names, SARIF rule ids), what semver means
    here, and that nothing is sent anywhere (no telemetry).
  • Action inputs sarif, junit and annotations; outputs gate-capped and report-path.
  • Pin the sample: --urls FILE scores exactly the given pages (up to 8, all on the scored site);
    --urls-from REPORT.json re-reads the pages an earlier report sampled. Sampling is the main source
    of run-to-run variance, so a before/after comparison on pinned pages shows what changed on the
    site. The report says sample_source: "pinned". Action input urls-from.
  • Per-check documentation links: every check in the JSON report carries doc_url, its definition
    in the rubric pinned to the release tag, so the link never drifts from the scorer. SARIF rules link
    there too, and --explain prints it under each check below full marks. Optional field in
    schema/report.v2.json.
  • Troubleshooting, in English and Chinese: score movement, checks shown
    as —, the gate cap, exit code 2, proxies, redirects, the leaderboard's tool version, watch
    exit 3, Action annotation limits, CJK thresholds. The terminal report links to it.
  • Intervals on the leaderboard's medians. benchmark/stats.py computes a seeded bootstrap 95%
    interval for the overall median, each sector median and the Chinese-market gap; benchmark/README.md
    and both READMEs quote them: overall 56 (53–58), the gap 19 points (12–26). The hosted page shows
    medians without intervals until it is rebuilt and redeployed. Sectors with fewer than 10 scored sites (China / cloud and Professional services, 9
    each) no longer print a median; validate.py counts scored sites, not listed ones.
    benchmark/run.py --readme-only regenerates the README without re-scoring.
  • The leaderboard can move host: GEO_SITE_BASE sets the base URL for canonical links, og:image,
    hreflang, llms.txt and the sitemap, and scripts/check_live.py compares a hosted page with the
    committed data (today's Pages build still shows 229 sites against 314).
  • How far to trust the CLI, measured. benchmark/validity.py re-scores the five published hand
    audits on the same pages and compares every check tier for tier: 79% of 86 check pairs agree (95%
    for checks read by rule, 64% for heuristics), and the CLI reads 3.6 points lower on average, from
    10 lower to 6 higher. Results, caveats and every disagreement: benchmark/VALIDITY.md (EN + zh).
    This replaces the unmeasured "typically 5 to 15 points lower", which was wrong for two sites in five.
  • Evidence basis for every check. rubric/v1.1.json gains an evidence object per check
    (basis, sources, last verified), published as rubric/evidence-v1.1.md (EN + zh). Metadata only:
    no weight, tier or score changed.
  • MCP for agents, not only for people. score_site, report, diff and list_runs return
    structuredContent (the full report, summary or diff as JSON) beside the text, and declare an
    outputSchema, so an agent no longer parses Markdown to recover a score. score_site takes urls
    (pin the pages, same checks as --urls) and sample. The rubric is served as the resource
    geo-score://rubric/v1.1, built from the scorer's own tables. The server instructions and the
    score_site description say that «site text: …» spans are data from the audited site.
  • watch diff --fail-on-drop: exit 4 when a verdict is change with a negative change. The MCP
    diff tool takes format: "json". guide/watch-methodology.md
    (EN + zh) gives the formulas and a worked example.
  • One gate: scripts/check.sh. Structure, ruff, the three test suites on Python 3.8 and the newest
    available, and an emulation of the GitHub Action against a fixture site (scripts/test_action.py);
    --live adds example.com. .githooks/pre-push runs it; .github/workflows/validate.yml runs the same
    script. Hosted Actions have not run on this organisation so far, so CONTRIBUTING.md no longer says
    CI checks anything: maintainers run the script before every merge.
  • AGENTS.md: the repository's invariants, each mapped to the check that enforces it, for
    contributors and coding agents (CLAUDE.md points to it).
  • Issue forms ask for the tool version and the JSON report; a new form covers level 2 and 3 engine
    problems and warns never to paste a key; security reports are routed to SECURITY.md.
  • Offline end-to-end tests. cli/test_geo_score_e2e.py serves twelve fixture sites from
    127.0.0.1 (a good site, a JavaScript shell, robots blocks, a firewall that refuses AI crawlers, a
    Chinese site, a site under a path, a hostile site that tries to address an agent, and more) and
    checks every check's tier against each site's expected.json, plus a golden report. A socket
    guard fails the run if anything leaves the machine; the suite takes under a second.
  • Tests that the CLI's embedded rubric (check ids, points, every tier's points and wording, bands,
    caps, bonus) equals rubric/v1.1.json.

Changed

  • watch diff is paired, clustered and corrected for multiple comparisons. It compares the
    questions both runs answered, with an exact paired test per row (McNemar when every question has
    one answer per run, otherwise an exact sign-flip test on each question's net change) and Holm's
    correction across the engine rows and the all-engines row. The verdict is change only when the
    adjusted p < 0.05, and too few questions to tell below 6 shared questions. 1.2.0 used an unpaired
    two-proportion test on answers that are not independent, which called changes too readily. New
    row fields: p_paired, p_adjusted, n_pairs, n_questions, n_asks, before_paired,
    after_paired, delta_pp, delta_ci. before, after, delta and p_value keep their 1.2.0
    unpaired meaning; no verdict uses them.
  • Intervals account for repeated asks. Every rate block gains cited_ci_clustered (a bootstrap
    over questions, 2,000 seeded resamples, never narrower than Wilson) and n_questions; reports quote
    it whenever a question has more than one answer. cited_ci (Wilson) is unchanged.
  • Headline rates leave out questions that name the brand. Questions are flagged at run time
    (brand_in_query, competitor_in_query) with the matcher used on answers; the headline covers the
    other questions and branded ones get their own line and n. An optional branded column in
    queries.csv (auto, yes, no) overrides the flag, watch check warns when more than 30% of the
    questions name the brand, and the CSV report gains a trailing branded column.
  • The Action scores the site once. It used to fetch and score twice (once for JSON, once for
    the log), which doubled the time and could print two different scores. The job summary, the
    annotations and every file now come from one run.
  • The Action annotates checks at zero: gates as errors, other checks as warnings (GitHub shows up
    to 10 of each; the rest are in the job summary).
  • geo-score watch --help names the command you typed (geo-score after an install).
  • --compare with three or more sites no longer says "Both scored".
  • When a site redirects to another domain, the full report says so and the JSON carries landed
    (optional), the URL that was scored.

Fixed

  • Scoring fixes the offline fixtures found. Each moves a score only where the old reading
    contradicted the rubric:
    • g.robots: several User-agent lines share the rules that follow them (RFC 9309 2.1). The rules
      were attached to the last line only, so a group listing GPTBot, ClaudeBot and CCBot over
      Disallow: / read as blocking only CCBot. Down, to 0 and the gate cap, where grouped rules block
      retrieval crawlers; up (3 to 5) where a grouped Allow now counts as an explicit permission.
    • p1.llms-txt, b.llms-full, b.ai-txt, p1.sitemap: a single-page app that answers every path
      ...
Read more

v1.2.0 — levels 2 and 3: does AI actually cite you?

Choose a tag to compare

@vortwang vortwang released this 26 Sep 17:34

Added — levels 2 and 3: does AI actually cite you?

geo-watch 0.1.0 is merged into this repository. One tool now measures both halves of AI
visibility: whether engines can cite a site (the score, free) and whether they do
(with the user's own API keys). The citation half is never summed into the score, as the
rubric has always required.

  • Level 2, --ask "question" (repeatable, up to 5): after the readiness report, asks
    ChatGPT, Perplexity, Gemini and Claude the question through their search-enabled APIs,
    one ask per engine that has a key, and fills the report's citation object. Without a key
    the block says not measured and the score is unchanged.
  • Level 3, geo_score.py watch: a fixed question list asked every week; citation rate per
    engine with 95% Wilson intervals, share of voice against competitors, the third-party
    sources the engines cite, the searches they ran, a cost ledger with caps (max_calls
    is exact; budget_usd can be exceeded by at most one call per engine), and diff with a
    two-proportion test. Runs follow schema/watch.v1.json.
  • geo_score.py mcp: a dependency-free stdio MCP server with score_site (level 1) and
    ask, run, list_runs, report, diff (levels 2 and 3).
  • schema/report.v2.json: the reserved citation object gains optional channel,
    answered, mentioned_count, per_engine, top_sources and cost_usd. Additive; every
    1.1.x report still validates.
  • examples/ci/watch-weekly.yml: the weekly schedule.

Levels 2 and 3 need cli/geo_watch.py from the same release next to cli/geo_score.py. The
one-line curl | python3 install still runs level 1 on its own. Scores are unchanged from 1.1.2.

For geo-watch 0.1.0 users: rename geo-watch.json to geo-score-watch.json and run
python3 cli/geo_score.py watch … instead of geo_watch.py …. Runs now live in
.geo-score/watch/runs/ with format geo-score/watch.v1; 0.1.0 run files are not read. The CSV
export renames competitors_cited to other_domains_cited and drops the always-empty
wrong_fact column.

Full changelog: CHANGELOG.md · 中文:CHANGELOG.zh-CN.md

v1.1.2 — a network that cannot reach something is not evidence it is missing

Choose a tag to compare

@vortwang vortwang released this 26 Sep 17:34

Changed — scores move

A site scored 88 and then 95 on two consecutive runs from one machine, and another
scored 71, 40, 71 — both outside the ±5 this repository promises. Three causes, one
principle: a network that cannot reach something is not evidence that it is missing.

  • p3.sameas counted every non-200 as a dead profile. Timeouts, resets and 451 (LinkedIn,
    X, YouTube, GitHub and Wikipedia are unreachable from mainland China) and bot walls
    (401 / 403 / 429 / 999 from Crunchbase, Zhihu, Facebook, LinkedIn) are now left out;
    only 404 / 410 count as gone. When no link can be checked, the check leaves the
    denominator instead of scoring 0. The same 12 profiles resolved 1 of 12 from a
    mainland network and 9 of 12 from overseas.
  • p3.knowledge-graph scored 0 when some Wikidata / Wikipedia lookups timed out and
    the rest found nothing. Absence now needs every lookup to answer; a partial answer is
    unobservable. (All lookups failing was already unobservable.)
  • Sampling fetched the sitemap once. One dropped request sent discovery to homepage
    links, a different eight pages were scored, and the middle run fell under the g.ssr
    gate (capped at 40). The homepage, robots.txt, sitemap, section indexes and sampled
    pages are now retried on timeouts, 429 and 5xx. Reports carry sample_source.

After the change, three runs each: jianruntech.com 95 / 95 / 95 (direct) and 93 / 93 / 95
(overseas); stripe.com 75 / 75 / 75 (direct) and 70 / 72 / 70 (overseas, concurrent).
The published benchmark was measured on 2026-09-09 with the earlier logic and has not been
re-run.
A site whose sameAs or knowledge-graph lookups hit an unreachable host can score a
few points higher today than its leaderboard row. Reports now say geo-score-cli/1.1.2, so the
two can be told apart.

Fixed

  • gzip bomb. The 4 MB read cap applied to the compressed stream, and gzip.decompress
    has no output limit, so a few MB of crafted gzip could inflate to gigabytes and kill the
    process — on a GitHub Action runner, an out-of-memory crash instead of a score. Inflation
    is now capped at 16 MB; real pages never get near it. No score changes.